Santhosh Kumar J
A software developer turned platform engineer, now building AI systems — 14+ years from embedded firmware and backend services to cloud-native platforms on AWS, Azure and GCP serving 10M+ users.

By the numbers
14+
Years of experience
50+
Production Kubernetes clusters managed
10M+
Users served by our platforms
99.9%
Availability sustained
Blogs
View allAgent memory and context engineering in Postgres (3/3)
An agent whose pod can vanish mid-investigation and whose workflow exits after thirty minutes idle — so every message it has ever seen is a row, in one table with a raw column and a compacted one, with tool output reduced before it reaches the context. Plus one incident thread walked row by row, from 42,709 tokens of telemetry down to 1,889.
Embedding and vector search for incident recall (2/3)
Step 6 of a first-level investigation — "have we seen this before?" — is the one a human always loses, because Slack search is keyword search. How pgvector makes it work: rendering alerts and resolved threads into the same shape so similarity means something, a cosine HNSW index and the filter trap hiding inside it, and a Kafka pipeline that re-embeds whole threads rather than messages.
Incident agents with LangGraph, LangChain and MCP (1/3)
How the investigation itself is built and run — MCP servers as the only door to Prometheus, Loki, Tempo and pod events; a LangGraph state graph running the telemetry sweep and the recall search in parallel; a ReAct loop confined to a single node; an LLM judge deciding whether anything gets posted; why none of it can live inside workflow code; and KEDA scaling the workers to zero between incidents.
GitHub Contributions
59 in the last yearImpact
FinOps governance
Rightsizing, autoscaling, Reserved Instance planning and Karpenter-based Spot optimisation across customer platforms.
35%
cloud costs cut
GitOps delivery
GitOps-driven deployment pipelines built with Argo CD, GitHub Actions, Jenkins and Azure DevOps.
80%
less deploy effort
Infrastructure as Code
Terraform, Kustomize and CloudFormation standards rolled out within 2 months — eliminating configuration drift.
70%
less provisioning
Self-service platform
Developer environments orchestrated in shared Kubernetes clusters with Argo CD and GitHub Actions.
20+
dev environments
Observability adoption
Self-hosted LGTM stack ingesting 20 GB of logs a day and 10M metric series — at 50% lower cost than managed tooling.
15+
teams onboarded
DevSecOps programmes
Vulnerability remediation and CIS-aligned security baselines across 10+ customer environments.
50+
CVEs resolved