Senior Observability Engineer

MX

Full Time5+ yearsPosted about 1 month ago

Let the right jobs find you

Get personalised suggestions from verified company career pages, matched to your role, location, level, and skills.

Overview

Position Type

Full Time

Experience

5+ years

Job Description

Job Duties:

  • Build and operate an observability control plane: automate baseline monitors, dashboards, and tagging standards through the Datadog API and Terraform.
  • After significant incidents, produce detection and dashboard gap packs grounded in Datadog and MX investigation patterns, with queries ready to apply.
  • Define what "good" looks like for a Ruby, Go, or Java service on Datadog (tags, golden signals, alert quality, dashboard contracts), then audit services against that standard and accept or reject readiness.
  • Validate, don't own. Service owners keep their alerts and dashboards; you confirm they are complete and correct, then move on. Escalate to engineering managers when coverage fails or an owner is missing.
  • Own the monthly observability and service-catalog health report: departed owners, stale dashboards, services with no monitors, SLO gaps, and coverage trends.
  • Run maturity assessments (baseline through SLO, launch-ready, self-serve) and track them over time.
  • Tune alerting toward zero false SEV1/2 pages and actionable SEV3/4 alerts, and coach teams on Datadog cost and cardinality.
  • Build self-serve onboarding so new services get baseline observability on day one, without a multi-week embed.
  • Share the team pager. Rotate on the shared IR & Observability on-call, triage and investigate live incidents with Datadog and MX investigation patterns, and take Incident Commander or supporting technical roles as the incident needs.
  • After incidents, close the detection loop (gap packs, new monitors, dashboards) so the pager gets quieter over time.
  • Run high-value launch and production-readiness reviews as a checkpoint, not a permanent staffing model.

Basic Requirements:

  • BS in Computer Science or equivalent experience
  • 5+ years running production observability, SRE, or DevOps. Datadog preferred; strong Grafana/Prometheus, Splunk, or New Relic experience counts if you can ramp on Datadog fast.
  • Automation-first engineering in Python, Bash, Go, and/or Terraform, plus Kubernetes proficiency. You encode monitoring standards as code rather than clicking the UI.
  • AI- and workflow-literate. You've used or built scripted and AI-assisted workflows to scale reviews, audits, and docs.
  • Alerting and SLO strategy: burn-rate and error-budget thinking, with a track record of cutting alert fatigue on evidence.
  • Distributed-systems debugging across microservices: latency, connection pools, queues, and cascading failure on Kubernetes and bare metal, with NATS, RabbitMQ, Postgres, and Redis.
  • Shared on-call, Incident Commander-capable. You've run or supported incident bridges and written postmortems, and you'll take shifts on the shared IR & Observability rotation.

Preferred Requirements:

  • Fintech experience with MX-like architectures
  • Google SRE practices: toil elimination, incident management, automation for self-healing
  • Cross-functional influence without authority. You've improved teams that don't report to you.
  • Governance and reporting: you can produce a monthly health and compliance report leadership reads (orphans, stale entries, gaps, trends).
  • OpenTelemetry instrumentation
  • Datadog cost optimization at scale (cardinality, log indexing, sampling)
  • Incident response platforms (incident.io, PagerDuty, OpsGenie); prior formal Incident Commander experience
  • Golang and Ruby on Rails (the MX stack)

What Success Looks Like:

  • By six months, you're a full participant on the shared on-call rotation and a capable Incident Commander on live SEVs, teams you've engaged have alerts and dashboards that answer "what's broken and where do I look?", and the monthly health report runs largely on its own. By twelve months, incidents get caught earlier because of instrumentation the loop added, new services get baseline observability from a self-serve template on day one, and no team depends on a shepherd for day-one coverage.

Required Skills

DatadogGrafanaPrometheusSplunkNew RelicPythonBashGoTerraformKubernetes

About the Company

MX

Chennai, India

Share This Job