Summary
Eight years of reliability work for payment and marketplace platforms: SLOs and error budgets, Kubernetes at scale, OpenTelemetry end to end, incident command and the engineering that follows. 99.99 % over the last eighteen months, with the receipts.
Projects
- Merged
slo-platform #1849
SLO definitions as code, burn-rate alerts, error-budget reports to every team lead on Monday morning. 210 SLOs, 140 services.
- Merged
incident-bot #1867
Opens the channel, pages the right people, keeps the timeline, drafts the postmortem. Nobody writes the timeline by hand any more.
- Merged
Chaos days #1598
Quarterly game days with real failure injection in production-like environments; 31 findings fixed so far.
Skills
Reliability
- SLOs and error budgets
- Incident command
- Capacity planning
- Chaos engineering
Observability
- Prometheus / Thanos
- Grafana
- OpenTelemetry
- Loki / Tempo
Platform
- Kubernetes
- Kafka
- PostgreSQL
- Terraform
Languages
- Go
- Python
- PromQL
Experience
Staff Site Reliability Engineer
Paylane · Bengaluru
Reliability lead for a payments platform; team of six; SLO platform, observability, incident process.
- 99.99 % availability over the last eighteen months against a 99.95 % SLO
- Mean time to recovery 54 → 12 minutes with alert routing on burn rate and runbook automation
- OpenTelemetry across 140 services; 60 % of dashboards retired, the rest used daily
Site Reliability Engineer
Marketo Bazaar · Bengaluru
Reliability for a marketplace with festival-season peaks of 30× normal traffic.
- Ran the first chaos days; found and fixed nine single points of failure before the sale season
- Kubernetes capacity model that cut over-provisioning 40 %
Backend Engineer
Finly · Bengaluru
Services for a lending product; became the person who got paged, then the person who fixed why.
Achievements
- feat: 99.99 % availability, eighteen months running
- feat: MTTR 54 → 12 minutes
- fix: Nine single points of failure found and fixed before the sale season
About
I lead reliability at Paylane, a payments platform doing 2,000 transactions a second at peak. My team of six owns the SLO platform, the observability stack, the incident process and the chaos days that keep everyone honest.
I believe the best incident is the one the error budget predicted a week earlier.
Publications
- Burn-rate alerts that people trustNov 2024 · Personal blog
- SREcon Asia — Retiring 60 % of our dashboardsJun 2025 · Talk
Certifications
- Certified Kubernetes AdministratorApr 2021 · CNCF
- Google Professional Cloud ArchitectJul 2023 · Google Cloud
- Prometheus Certified AssociateJan 2024 · CNCF
Education
B.Tech, Computer Science and Engineering
PES University
Testimonials
Divya turned our incidents from panic into procedure, and then into fewer incidents.
Karthik Iyer — CTO, Paylane
Languages
- EnglishFluent
- MalayalamNative
- KannadaConversational
Interests
- Marathons
- Board games
- Pottery