AI-Native Reliability

If your system works for 400 users and now needs to work for 4 million, that’s where I come in.

VP of Site Reliability Engineering. I build AI-native reliability systems — instrumentation, diagnosis, and response that happen before a human is paged.

Read the work →

96%
MTTR reduction — 1,400h to 60h JPMC
3M
concurrent users, 0 P0 on launch day WBD Max
2PB+
migrated with zero downtime JPMC
100%
uptime through the Olympics 2024 peak WBD

Work

Overview — deep-dives arrive in a later phase

AEGIS

Deep-dive

Incidents still wait on a human to read the alert, write the runbook, and dispatch the fix.

Role
Architect & builder
Timeframe
2024–present

An agentic pipeline: a Strategist decomposes the alert, composes a dynamic runbook, and dispatches sub-agents across observability and ticketing surfaces via MCP. Human judgment stays at the decision boundary.

Pattern validated in a production enterprise environment.

independent build — available on request

O11y Skill Suite

Deep-dive

Observability standards usually live in documents that no runtime enforces.

Role
Creator
Timeframe
2024–present

A skill library and MCP servers that encode OTel instrumentation and o11y governance — standards as running code, not review checklists — so the practice scales uniformly across hybrid-cloud platforms. Outcome quantified in the Phase 3 write-up.

standalone demonstration project

Max at Scale

Launch-day traffic across three regions is not a load test you can rerun.

Role
SRE lead
Timeframe
2022–2024

Self-service load testing validated 3M concurrent streams. 0 P0 incidents on global launch day; 100% uptime through the Olympics 2024 peak.

employer work — public metrics

JPMC Reliability

SLO governance across 12 critical services and 100+ APIs.

Role
VP, Site Reliability Engineering
Timeframe
2024–present

MTTR from 1,400 hours to 60. P1/P2 incidents near zero.

employer work — public metrics

Mindtree Observability

Monitoring lived as a hand-configured service-desk function, stale the day after it shipped.

Role
New Relic SME
Timeframe
2013–2022

95% infrastructure coverage; dashboard-as-code cut deployment time 75%.

employer work — public metrics

Contact

One channel

For a Head-of-SRE conversation, an AEGIS walkthrough, or a good argument about autonomous systems.

jeyadev.narayanan@gmail.com