All systems nominal — achievement unlocked

visitor@shubhamkumar:~ ● online

AI Platform OperationsKubernetes Platform EngineeringCI/CD AutomationObservability & Incident ResponseMulti-Cloud Infrastructure
Remote-friendlyPlatform-first mindsetBuilt for production
Shubham Kumar
TENURE

0+

Years building in production

UPTIME_IMPACT

0%

Fewer production outages

VELOCITY

0%

Faster release workflows

HOURS_RECLAIMED

0+

Hours saved each month

Kubernetes
AWS
GCP
Azure
AWS Bedrock
Terraform
Karpenter
CI/CD
RAG Pipelines
Observability
FinOps
Linux
DevOps
Zero Trust / Keycloak
Platform Engineering
Automation
Incident Response

About

The engineer behind the platform work

Most of my work sits at the intersection of cloud infrastructure, automation, and reliability. I enjoy building systems that make delivery smoother and production less stressful.

Cloud Platform Engineer | DevOps, Kubernetes & Automation

I started where most infrastructure careers do: cloud fundamentals. Kubernetes operations, infrastructure as code, CI/CD automation, and multi-cloud reliability work across AWS, Azure, and GCP. That foundation is still how I think about systems. Observable, automated, and boring in the best way. What's changed is where I'm pointing that foundation. Over the past year, as a Cloud Platform Engineer, I've been increasingly focused on building and operating production AI tooling on AWS Bedrock with Claude models, including a RAG-based auto-triage agent and an AI-assisted incident RCA tool, both deliberately human-gated rather than fully autonomous. I pair that with FinOps cost governance and Zero Trust identity (Keycloak, OIDC/OAuth2). It's the unglamorous platform work that has to be solid before you can trust an AI system to touch production. That's the direction I'm building toward: applying DevOps discipline (guardrails, observability, human-in-the-loop review) to systems that increasingly involve LLMs in the operational path.

Experience

Teams and systems I've worked on

From embedded environments to cloud platforms, most of my roles have been about improving reliability, automating the rough edges, and helping teams ship with more confidence.

SingleStore
SingleStorestatus: active

Cloud Platform Engineer

Remote

Feb 2026PresentRemoteAI Platform
  • Designed and built ATLAS, an internal operational-intelligence platform (Airflow, SingleStore, Next.js) used daily by Support, Engineering, and Leadership to track SLA risk and recurring issues.
  • Built production AI tooling on AWS Bedrock (Claude 3/3.5): an AI-assisted incident RCA tool on Grafana MCP and a guardrailed RAG auto-triage agent, both human-gated by design.
  • Implemented GPU-backed autoscaling on EKS using Karpenter, cutting ML infrastructure cost by 35%+ (utilization ~25% to 65%), measured via Kubecost against real AWS billing.
  • Own the production Keycloak identity platform (Zero Trust, OIDC/OAuth2) serving ~150 daily internal users.
  • Contribute Go backend code to an internal multi-cloud cost-governance (FinOps) platform.

0%+

Cost Reduction

~0

Daily Users

Senior DevOps Engineer

Bengaluru, Karnataka, India

Aug 2025Feb 2026BengaluruPromotion
  • Led an uptime initiative that raised platform availability from 97.8% to 99.95%.
  • Cut incident resolution time by 45%+ through centralized observability and standardized runbooks.
  • Mentored junior engineers and set incident-response and IaC standards adopted across every team on the shared platform.

97.80%

Platform Uptime

0%+

Incident Resolution

DevOps Engineer

Bengaluru, Karnataka, India

Oct 2023Jul 2025BengaluruKubernetesEmbedded Systems
  • Developed and deployed automation across a fleet of 8,000+ embedded IFE (in-flight entertainment) devices. This included firmware rollout pipelines that cut release time by 40%, plus telemetry-based PMIC monitoring with secure LTE-based diagnostics that reduced MTTR by 35%.
  • Built DISCO, an internal Python-based tool for processing onboard infotainment box log data at scale, pulling and parsing logs from AWS S3 for fleet-wide diagnostics.
  • Operated and optimized AWS and Azure Kubernetes environments for production workloads; automated infrastructure changes with Terraform and CI-driven workflows.
  • Implemented monitoring and alerting improvements that reduced production outages by 40%.

0K+

Embedded Devices

0%

Firmware Release Time Cut

Innoitus

Innoitus

Site Reliability Engineer

Bengaluru, Karnataka, India

Jun 2023Sep 2023BengaluruMonitoring
  • Improved observability and alert quality through custom tooling and hands-on monitoring improvements.
  • Reduced critical incident frequency by 35% through proactive monitoring and reliability practices.
  • Improved incident response times by 30% with better alerting and on-call workflows.

0%

Fewer Incidents

0%

Faster Response

Quality Analyst

Bengaluru, Karnataka, India

Oct 2021Jun 2023BengaluruAWS
  • Built Jenkins pipelines integrating Prometheus and Grafana dashboards for better pipeline and environment visibility.
  • Managed AWS-based environments with a focus on scalability, uptime, and dependable delivery workflows.
  • Administered Kubernetes workloads with resource optimization across QA and production-adjacent systems.
Extreme Soft Management

Extreme Soft Management

Site Reliability Engineer

Ranchi, Jharkhand, India

Apr 2019Aug 2021RanchiGCP
  • Operated and maintained production infrastructure on Google Cloud Platform (GCP), introducing automation for repetitive operational tasks.
  • Led a year-long GCP-to-AWS migration, modernizing the deployment stack end-to-end.
  • Automated workflows that saved 80+ engineering hours per month across recurring processes.

0%

Fewer Environment Issues

0+

Hours Saved / Month

Tech Stack

What I reach for most

A practical mix of tools I use across cloud infrastructure, automation, observability, platform engineering, and day-to-day problem solving.

What I Do

Where I add the most value

This is the kind of work I usually take ownership of when a team wants faster delivery, better visibility, and systems that hold up in production.

01

Platform Engineering & Multi-Cloud Architecture

Designing secure, scalable cloud platforms across AWS, Azure, and GCP using Kubernetes, infrastructure as code, and CI/CD, with a strong focus on maintainability.

02

AI Platform Operations & FinOps

Operating production AI/LLM tooling on AWS Bedrock (Claude models): RAG pipelines and AI-assisted incident response (human-gated by design), alongside cost governance and FinOps.

03

Reliability, Zero Trust & Incident Response

Improving uptime and operational readiness through centralized observability and runbooks, backed by Zero Trust identity (Keycloak, OIDC/OAuth2).

Selected Work

A few things I've built and improved

These projects reflect the kind of work I enjoy most: strengthening platforms, simplifying operations, and building reliable workflows around real systems.

What people I've worked with say

A few words from teammates and engineering leaders who've seen my work up close.

I have had the opportunity to work with Shubham and closely observe his technical depth, execution ownership, and problem-solving mindset. Shubham brings a rare combination of strong Linux fundamentals, cloud-native expertise, and operational maturity. He approaches DevOps not just as tooling, but as an engineering discipline focused on reliability, scalability, and long-term maintainability. During our collaboration, I found him to be proactive, detail-oriented, and calm under pressure. He takes complete ownership of complex infrastructure challenges and drives them to closure without noise. His ability to bridge embedded systems, cloud platforms, and automation pipelines makes him particularly valuable in IoT and distributed environments. Beyond technical skills, Shubham is dependable, collaborative, and always willing to go the extra mile for the team. I strongly recommend him for any role that demands both technical excellence and execution discipline.

Yash Anand
Yash Anand

Director of Technology · AirFi · Mar 2026

Get In Touch

If you'd like to work together

I'm always open to a good conversation around platform engineering, DevOps, cloud infrastructure, or interesting production problems.

Have something interesting in mind?

Open to conversations about AI platform operations, DevOps automation, FinOps, and cloud reliability at scale.