Skip to content

Open to work — AI Engineer, Software Developer & Platform Engineer roles · full-time or consulting · open to relocation worldwide

Santhosh Kumar J

AI Engineer · Software Developer · Platform EngineerChennai, India

A software developer turned platform engineer, now building AI systems — 14+ years from embedded firmware and backend services to cloud-native platforms on AWS, Azure and GCP serving 10M+ users.

Portrait of Santhosh Kumar J, AI Engineer · Software Developer · Platform Engineer

By the numbers

14+

Years of experience

50+

Production Kubernetes clusters managed

10M+

Users served by our platforms

99.9%

Availability sustained

Agent memory and context engineering in Postgres (3/3)

An agent whose pod can vanish mid-investigation and whose workflow exits after thirty minutes idle — so every message it has ever seen is a row, in one table with a raw column and a compacted one, with tool output reduced before it reaches the context. Plus one incident thread walked row by row, from 42,709 tokens of telemetry down to 1,889.

AISREPostgreSQL+1

Embedding and vector search for incident recall (2/3)

Step 6 of a first-level investigation — "have we seen this before?" — is the one a human always loses, because Slack search is keyword search. How pgvector makes it work: rendering alerts and resolved threads into the same shape so similarity means something, a cosine HNSW index and the filter trap hiding inside it, and a Kafka pipeline that re-embeds whole threads rather than messages.

AIPostgreSQLRAG+1

Incident agents with LangGraph, LangChain and MCP (1/3)

How the investigation itself is built and run — MCP servers as the only door to Prometheus, Loki, Tempo and pod events; a LangGraph state graph running the telemetry sweep and the recall search in parallel; a ReAct loop confined to a single node; an LLM judge deciding whether anything gets posted; why none of it can live inside workflow code; and KEDA scaling the workers to zero between incidents.

AILangChainMCP+1

GitHub Contributions

59 in the last year

Impact

FinOps governance

Rightsizing, autoscaling, Reserved Instance planning and Karpenter-based Spot optimisation across customer platforms.

35%

cloud costs cut

GitOps delivery

GitOps-driven deployment pipelines built with Argo CD, GitHub Actions, Jenkins and Azure DevOps.

80%

less deploy effort

Infrastructure as Code

Terraform, Kustomize and CloudFormation standards rolled out within 2 months — eliminating configuration drift.

70%

less provisioning

Self-service platform

Developer environments orchestrated in shared Kubernetes clusters with Argo CD and GitHub Actions.

20+

dev environments

Observability adoption

Self-hosted LGTM stack ingesting 20 GB of logs a day and 10M metric series — at 50% lower cost than managed tooling.

15+

teams onboarded

DevSecOps programmes

Vulnerability remediation and CIS-aligned security baselines across 10+ customer environments.

50+

CVEs resolved

Skills & Tools

KubernetesDockerHelmAmazon Web Services (AWS)Google Cloud Platform (GCP)Microsoft AzureCloudflareTerraformAWS CloudFormationGitOpsArgo CDGitHub ActionsJenkinsAzure DevOpsPrometheusGrafanaLokiTempoOpenTelemetrySigNozNew RelicTypeScriptNode.jsReactNext.jsPostgreSQLMySQLMongoDBRedisElasticsearchNeo4jLinuxShell ScriptingGit