Internal developer platforms, GitOps delivery, Kubernetes at multi-team scale, and the observability that keeps it all reliable.
Most recently at Aldi Süd: led an org-wide observability migration from New Relic to Dynatrace across multiple engineering departments, standardised on OpenTelemetry with monitoring-as-code in Terraform, and cut MTTR and alert noise ~30%.
📍 Berlin, Germany · Permanent residence (Niederlassungserlaubnis) — no visa sponsorship required
Portfolio & architecture case studies → · LinkedIn
Each links to a case study with the architecture decisions and the trade-offs I rejected.
OpenTelemetry & LGTM Platform — multi-cluster observability on EKS: OTel agent→gateway, tail sampling, Mimir + Loki + Tempo on S3
↳ github.com/ok-karthik/opentelemetry-platform-on-eks
IDP & GitOps Reference Architecture — Go scaffolder CLI, versioned Terraform modules, Argo CD ApplicationSets, multi-tenant delivery
↳ github.com/ok-karthik/internal-developer-platform
Enterprise AWS Infrastructure — multi-environment AWS in Terragrunt/Terraform: DRY module hierarchy, OPA policy gates, self-healing CI
↳ github.com/ok-karthik/enterprise-aws-infrastructure-terragrunt
AI Infrastructure on EKS — GPU workloads: Karpenter Spot autoscaling, NVIDIA GPU Operator, time-slicing, DCGM metrics
↳ github.com/ok-karthik/ai-infrastructure-on-eks
FinOps Kubernetes Operator — scales non-production workloads to zero on a namespace annotation, restores them on schedule
↳ github.com/ok-karthik/finops-k8s-operator
Also: observable-inference-gateway — LLM gateway with end-to-end OTel tracing · job-market-radar — ETL + analytics measuring what job ads actually require, not just mention
Open to Senior/Staff Platform Engineering and SRE roles in Germany


