visitor@shubhamkumar:~ ● online

0+
Years building in production
0%
Fewer production outages
0%
Faster release workflows
0+
Hours saved each month
About
The engineer behind the platform work
Most of my work sits at the intersection of cloud infrastructure, automation, and reliability. I enjoy building systems that make delivery smoother and production less stressful.
Cloud Platform Engineer | DevOps, Kubernetes & Automation
I started where most infrastructure careers do: cloud fundamentals. Kubernetes operations, infrastructure as code, CI/CD automation, and multi-cloud reliability work across AWS, Azure, and GCP. That foundation is still how I think about systems. Observable, automated, and boring in the best way. What's changed is where I'm pointing that foundation. Over the past year, as a Cloud Platform Engineer, I've been increasingly focused on building and operating production AI tooling on AWS Bedrock with Claude models, including a RAG-based auto-triage agent and an AI-assisted incident RCA tool, both deliberately human-gated rather than fully autonomous. I pair that with FinOps cost governance and Zero Trust identity (Keycloak, OIDC/OAuth2). It's the unglamorous platform work that has to be solid before you can trust an AI system to touch production. That's the direction I'm building toward: applying DevOps discipline (guardrails, observability, human-in-the-loop review) to systems that increasingly involve LLMs in the operational path.
Contact Details
Experience
Teams and systems I've worked on
From embedded environments to cloud platforms, most of my roles have been about improving reliability, automating the rough edges, and helping teams ship with more confidence.
- Designed and built ATLAS, an internal operational-intelligence platform (Airflow, SingleStore, Next.js) used daily by Support, Engineering, and Leadership to track SLA risk and recurring issues.
- Built production AI tooling on AWS Bedrock (Claude 3/3.5): an AI-assisted incident RCA tool on Grafana MCP and a guardrailed RAG auto-triage agent, both human-gated by design.
- Implemented GPU-backed autoscaling on EKS using Karpenter, cutting ML infrastructure cost by 35%+ (utilization ~25% to 65%), measured via Kubecost against real AWS billing.
- Own the production Keycloak identity platform (Zero Trust, OIDC/OAuth2) serving ~150 daily internal users.
- Contribute Go backend code to an internal multi-cloud cost-governance (FinOps) platform.
0%+
Cost Reduction
~0
Daily Users
- Led an uptime initiative that raised platform availability from 97.8% to 99.95%.
- Cut incident resolution time by 45%+ through centralized observability and standardized runbooks.
- Mentored junior engineers and set incident-response and IaC standards adopted across every team on the shared platform.
97.80%
Platform Uptime
0%+
Incident Resolution
- Developed and deployed automation across a fleet of 8,000+ embedded IFE (in-flight entertainment) devices. This included firmware rollout pipelines that cut release time by 40%, plus telemetry-based PMIC monitoring with secure LTE-based diagnostics that reduced MTTR by 35%.
- Built DISCO, an internal Python-based tool for processing onboard infotainment box log data at scale, pulling and parsing logs from AWS S3 for fleet-wide diagnostics.
- Operated and optimized AWS and Azure Kubernetes environments for production workloads; automated infrastructure changes with Terraform and CI-driven workflows.
- Implemented monitoring and alerting improvements that reduced production outages by 40%.
0K+
Embedded Devices
0%
Firmware Release Time Cut
Innoitus
Site Reliability Engineer
Bengaluru, Karnataka, India
- Improved observability and alert quality through custom tooling and hands-on monitoring improvements.
- Reduced critical incident frequency by 35% through proactive monitoring and reliability practices.
- Improved incident response times by 30% with better alerting and on-call workflows.
0%
Fewer Incidents
0%
Faster Response
- Built Jenkins pipelines integrating Prometheus and Grafana dashboards for better pipeline and environment visibility.
- Managed AWS-based environments with a focus on scalability, uptime, and dependable delivery workflows.
- Administered Kubernetes workloads with resource optimization across QA and production-adjacent systems.
Extreme Soft Management
Site Reliability Engineer
Ranchi, Jharkhand, India
- Operated and maintained production infrastructure on Google Cloud Platform (GCP), introducing automation for repetitive operational tasks.
- Led a year-long GCP-to-AWS migration, modernizing the deployment stack end-to-end.
- Automated workflows that saved 80+ engineering hours per month across recurring processes.
0%
Fewer Environment Issues
0+
Hours Saved / Month
Tech Stack
What I reach for most
A practical mix of tools I use across cloud infrastructure, automation, observability, platform engineering, and day-to-day problem solving.
Languages
Frameworks & Runtime
Databases
Specializations
What I Do
Where I add the most value
This is the kind of work I usually take ownership of when a team wants faster delivery, better visibility, and systems that hold up in production.
01
Platform Engineering & Multi-Cloud Architecture
Designing secure, scalable cloud platforms across AWS, Azure, and GCP using Kubernetes, infrastructure as code, and CI/CD, with a strong focus on maintainability.
02
AI Platform Operations & FinOps
Operating production AI/LLM tooling on AWS Bedrock (Claude models): RAG pipelines and AI-assisted incident response (human-gated by design), alongside cost governance and FinOps.
03
Reliability, Zero Trust & Incident Response
Improving uptime and operational readiness through centralized observability and runbooks, backed by Zero Trust identity (Keycloak, OIDC/OAuth2).
Selected Work
A few things I've built and improved
These projects reflect the kind of work I enjoy most: strengthening platforms, simplifying operations, and building reliable workflows around real systems.
What people I've worked with say
A few words from teammates and engineering leaders who've seen my work up close.
I have had the opportunity to work with Shubham and closely observe his technical depth, execution ownership, and problem-solving mindset. Shubham brings a rare combination of strong Linux fundamentals, cloud-native expertise, and operational maturity. He approaches DevOps not just as tooling, but as an engineering discipline focused on reliability, scalability, and long-term maintainability. During our collaboration, I found him to be proactive, detail-oriented, and calm under pressure. He takes complete ownership of complex infrastructure challenges and drives them to closure without noise. His ability to bridge embedded systems, cloud platforms, and automation pipelines makes him particularly valuable in IoT and distributed environments. Beyond technical skills, Shubham is dependable, collaborative, and always willing to go the extra mile for the team. I strongly recommend him for any role that demands both technical excellence and execution discipline.

Director of Technology · AirFi · Mar 2026
Get In Touch
If you'd like to work together
I'm always open to a good conversation around platform engineering, DevOps, cloud infrastructure, or interesting production problems.
Have something interesting in mind?
Open to conversations about AI platform operations, DevOps automation, FinOps, and cloud reliability at scale.