● ● ●shubham@gulati: ~/portfoliossh · live
$ whoami

Shubham Gulati

Principal Software Engineer

I build large-scale distributed systems, AI infrastructure, and GPU compute platforms — from streaming fault detection with Flink and Kafka to cloud services operating at million-QPS scale.

$ help
shubham@gulati:~$
01 / about

Systems that hold up when scale becomes real.

Principal engineer focused on distributed systems, databases, cloud infrastructure, GenAI, and AI infrastructure. I own systems from architecture through production, balancing scalability, reliability, security, performance, and cost.

02 / experience

Production systems, end to end.

Oct 2025 — Present · Oracle Cloud Infrastructure

Principal Member of Technical Staff

AI Infrastructure · GPU Observability · Developer Tooling

  • Building GPU observability pipelines that process cluster telemetry from Object Storage through Apache Flink, detect actionable faults, and publish events to Apache Kafka.
  • Architected OCI GenAI control/data-plane services, NL-to-SQL workflows, Code Interpreter, agentic shell tooling, and MCP integrations.
Aug 2022 — Oct 2025 · Oracle Cloud Infrastructure

Senior Member of Technical Staff

Led Big Data Gateway architecture and production launch; designed capacity for ~1M QPS while reducing compute footprint from 85 to 32 nodes per region.

Feb 2020 — May 2022 · Amazon Web Services

Software Engineer — DynamoDB (JournalDB)

Built multi-tenancy, infrastructure automation, and cross-account VPC connectivity for distributed database workloads.

2015 — 2018 · Zomato / ShopClues

Software Engineer

Built production backend systems and microservices for high-traffic consumer and seller platforms.

03 / selected systems projects

Things I like building.

FLINK · KAFKA · GPU INFRASTRUCTURE

GPU Fault Detection Pipeline

Object Storage → Apache Flink → fault detection → Apache Kafka → downstream customer workflows.

PYTHON · OPENTELEMETRY

Durable Agent Execution Runtime

Fault-tolerant agent execution with idempotent retries, event-sourced durability, deterministic replay, concurrency control, and tracing.

PYTHON · LLM INFRASTRUCTURE

KV Cache Manager Simulator

PagedAttention-style KV cache allocator with copy-on-write, prefix sharing, eviction policies, and benchmarking.

04 / stack

Tools are implementation details. Systems thinking is the skill.

JavaPythonGoRustApache FlinkApache KafkaDynamoDBKubernetes / OKEKubeletDockerTerraformOpenTelemetryGPU ComputeGPU ObservabilityMCPNL-to-SQLAWSOCI
05 / contact

Interested in hard infrastructure problems?