Skip to content
passingrun #8main4 stagesBengaluru, India
Divya Menon

Staff Site Reliability Engineer · Paylane

Divya Menon

Site reliability engineer — SLOs, observability and incidents that end with a better system

Error budgets instead of arguments; dashboards people actually look at; postmortems without blame.

  1. PES UniversityB.Tech, Computer Science and EngineeringAug 2014 – May 2018 — source
  2. FinlyBackend EngineerJun 2018 – Sep 2019 — completed
  3. Marketo BazaarSite Reliability EngineerOct 2019 – Aug 2022 — completed
  4. PaylaneStaff Site Reliability EngineerSep 2022 – Present — current
  • 8years of experience
  • 3roles
  • 3projects
  • 3certifications

Summary

Eight years of reliability work for payment and marketplace platforms: SLOs and error budgets, Kubernetes at scale, OpenTelemetry end to end, incident command and the engineering that follows. 99.99 % over the last eighteen months, with the receipts.

Experience

  1. Staff Site Reliability Engineer

    Paylane · Bengaluru

    Reliability lead for a payments platform; team of six; SLO platform, observability, incident process.

    • 99.99 % availability over the last eighteen months against a 99.95 % SLO
    • Mean time to recovery 54 → 12 minutes with alert routing on burn rate and runbook automation
    • OpenTelemetry across 140 services; 60 % of dashboards retired, the rest used daily
  2. Site Reliability Engineer

    Marketo Bazaar · Bengaluru

    Reliability for a marketplace with festival-season peaks of 30× normal traffic.

    • Ran the first chaos days; found and fixed nine single points of failure before the sale season
    • Kubernetes capacity model that cut over-provisioning 40 %
  3. Backend Engineer

    Finly · Bengaluru

    Services for a lending product; became the person who got paged, then the person who fixed why.

Projects

  • slo-platform

    slo-platform

    Lead · 2023 – present

    SLO definitions as code, burn-rate alerts, error-budget reports to every team lead on Monday morning. 210 SLOs, 140 services.

    • SLOs
    • Prometheus
    • Go
  • incident-bot

    incident-bot

    Author · 2024

    Opens the channel, pages the right people, keeps the timeline, drafts the postmortem. Nobody writes the timeline by hand any more.

    • Go
    • Slack
    • open source
    View
  • Chaos days

    Chaos days

    Organiser · 2020 – present

    Quarterly game days with real failure injection in production-like environments; 31 findings fixed so far.

    • chaos engineering
    • resilience

Skills

Reliability

  • SLOs and error budgets
  • Incident command
  • Capacity planning
  • Chaos engineering

Observability

  • Prometheus / Thanos
  • Grafana
  • OpenTelemetry
  • Loki / Tempo

Platform

  • Kubernetes
  • Kafka
  • PostgreSQL
  • Terraform

Languages

  • Go
  • Python
  • PromQL

Certifications

  • Certified Kubernetes Administrator

    CNCF

  • Google Professional Cloud Architect

    Google Cloud

  • Prometheus Certified Associate

    CNCF

Achievements

  • 99.99 % availability, eighteen months running2025
  • MTTR 54 → 12 minutes2024
  • Nine single points of failure found and fixed before the sale season2020

About

I lead reliability at Paylane, a payments platform doing 2,000 transactions a second at peak. My team of six owns the SLO platform, the observability stack, the incident process and the chaos days that keep everyone honest.

I believe the best incident is the one the error budget predicted a week earlier.

Education

  1. B.Tech, Computer Science and Engineering

    PES University

Publications

Testimonials

  • Divya turned our incidents from panic into procedure, and then into fewer incidents.

    Karthik Iyer — CTO, Paylane

Languages

  • EnglishFluent
  • MalayalamNative
  • KannadaConversational

Interests

  • Marathons
  • Board games
  • Pottery

Contact

Your message goes straight to Divya Menon. No account needed.

Theme

Palette

Mode

This template on