Senior Platform & Cloud Engineer who has run large-scale Kubernetes on AWS — 150+ EKS clusters, ~1,000 nodes, 30K+ pods serving 40+ engineering teams — and built an LLM inference platform on 500+ Spot GPUs at ~60% lower compute cost. Now at Aivar Innovation — EKS, Flyte orchestration, and GPU infrastructure for ML on KhiladiPro.
Case studies — real problems at 150-cluster, 500-GPU scale, solved with platforms and automation rather than headcount. Each one is a deep dive with architecture and impact.
A Kubernetes platform that schedules AI batch workloads on interruptible GPU capacity — checkpointing, retries, and Karpenter elasticity.
A Terraform module that turns ~50 lines of config into a full ECS Fargate platform — ALB routing, blue/green deploys, a pipeline per service.
A 4-phase orchestrator that upgrades EKS clusters fleet-wide — concurrent, cross-account, with automated post-upgrade validation.
ApplicationSets + multi-source Helm + a per-cluster Git inventory: one commit rolls an addon out to any cluster in the fleet.
Async fleet monitoring — concurrent cluster polling, thousands of endpoint probes with aiohttp, Prometheus metrics, actionable reports.
Bulk modernization of CI/CD stacks fleet-wide — idempotent, audited, ring-based rollout, safe to re-run after any partial failure.
REST API + React UI that onboards teams end-to-end — Terraform generation, Git automation, notifications — plus multi-tenant app patterns.
Plan → act → critique until evidence converges. MCP tools, Langfuse-gated evals, RBAC-scoped reads, writes only as GitOps PRs.
Discovers every ingress and service, probes them concurrently with asyncio — because "Running" doesn't mean "up." Prometheus-native.
It's real — type help, kubectl get pods, helm install manan or sudo and see what happens.
Whether you're a startup laying foundations or an enterprise wrangling hundreds of accounts — the playbook is turning recurring problems into systems.
Reusable IaC modules and golden paths so product teams ship infrastructure in minutes, not sprint cycles.
Upgrades, addon management, and health monitoring designed for hundreds of clusters — not one pet cluster.
Blue/green and canary deployments, circuit-breaker rollbacks, zero-downtime releases as the default.
Multi-account AWS design, landing zones, migration paths, and cross-account automation that scales with your org.
Monitoring and health tooling that surfaces issues before your users (or your pager) do.
Anything a team does manually more than twice a quarter becomes an automated, audited workflow.
Seven deep-dives — the 4-layer GPU optimization stack, interruption-native Spot GPUs, fleet upgrade doctrine, inventory-driven GitOps, the grounded agent, and the kro talk.
↗ ◆ ventures · builder mindsetGlassMic — digitally transforming decorative & packaging glass manufacturing — and AI Smart Insole, an iPhone-scan-to-custom-insole product concept. Complete solutions, not just code.
I'm heads-down building at Aivar Innovation — not job hunting — but I'm always up for talking infrastructure, GPUs, and agents.