Summary
Eight years of reliability work for payment and marketplace platforms: SLOs and error budgets, Kubernetes at scale, OpenTelemetry end to end, incident command and the engineering that follows. 99.99 % over the last eighteen months, with the receipts.
Experience
Staff Site Reliability Engineer
Paylane · Bengaluru
Reliability lead for a payments platform; team of six; SLO platform, observability, incident process.
- 99.99 % availability over the last eighteen months against a 99.95 % SLO
- Mean time to recovery 54 → 12 minutes with alert routing on burn rate and runbook automation
- OpenTelemetry across 140 services; 60 % of dashboards retired, the rest used daily
Site Reliability Engineer
Marketo Bazaar · Bengaluru
Reliability for a marketplace with festival-season peaks of 30× normal traffic.
- Ran the first chaos days; found and fixed nine single points of failure before the sale season
- Kubernetes capacity model that cut over-provisioning 40 %
Backend Engineer
Finly · Bengaluru
Services for a lending product; became the person who got paged, then the person who fixed why.
Projects
slo-platform
Lead · 2023 – present
SLO definitions as code, burn-rate alerts, error-budget reports to every team lead on Monday morning. 210 SLOs, 140 services.
incident-bot
Author · 2024
Opens the channel, pages the right people, keeps the timeline, drafts the postmortem. Nobody writes the timeline by hand any more.
ViewChaos days
Organiser · 2020 – present
Quarterly game days with real failure injection in production-like environments; 31 findings fixed so far.
Skills
Reliability
- SLOs and error budgets
- Incident command
- Capacity planning
- Chaos engineering
Observability
- Prometheus / Thanos
- Grafana
- OpenTelemetry
- Loki / Tempo
Platform
- Kubernetes
- Kafka
- PostgreSQL
- Terraform
Languages
- Go
- Python
- PromQL
Certifications
Certified Kubernetes Administrator
CNCF
Google Professional Cloud Architect
Google Cloud
Prometheus Certified Associate
CNCF
Achievements
- 99.99 % availability, eighteen months running2025
- MTTR 54 → 12 minutes2024
- Nine single points of failure found and fixed before the sale season2020
About
I lead reliability at Paylane, a payments platform doing 2,000 transactions a second at peak. My team of six owns the SLO platform, the observability stack, the incident process and the chaos days that keep everyone honest.
I believe the best incident is the one the error budget predicted a week earlier.
Education
B.Tech, Computer Science and Engineering
PES University
Publications
- Burn-rate alerts that people trustNov 2024 · Personal blog
- SREcon Asia — Retiring 60 % of our dashboardsJun 2025 · Talk
Testimonials
Divya turned our incidents from panic into procedure, and then into fewer incidents.
Karthik Iyer — CTO, Paylane
Languages
- EnglishFluent
- MalayalamNative
- KannadaConversational
Interests
- Marathons
- Board games
- Pottery