Senior Observability Engineer
MX
Full Time5+ yearsPosted about 1 month ago
Let the right jobs find you
Get personalised suggestions from verified company career pages, matched to your role, location, level, and skills.
Overview
Position Type
Full Time
Experience
5+ years
Job Description
Job Duties:
- Build and operate an observability control plane: automate baseline monitors, dashboards, and tagging standards through the Datadog API and Terraform.
- After significant incidents, produce detection and dashboard gap packs grounded in Datadog and MX investigation patterns, with queries ready to apply.
- Define what "good" looks like for a Ruby, Go, or Java service on Datadog (tags, golden signals, alert quality, dashboard contracts), then audit services against that standard and accept or reject readiness.
- Validate, don't own. Service owners keep their alerts and dashboards; you confirm they are complete and correct, then move on. Escalate to engineering managers when coverage fails or an owner is missing.
- Own the monthly observability and service-catalog health report: departed owners, stale dashboards, services with no monitors, SLO gaps, and coverage trends.
- Run maturity assessments (baseline through SLO, launch-ready, self-serve) and track them over time.
- Tune alerting toward zero false SEV1/2 pages and actionable SEV3/4 alerts, and coach teams on Datadog cost and cardinality.
- Build self-serve onboarding so new services get baseline observability on day one, without a multi-week embed.
- Share the team pager. Rotate on the shared IR & Observability on-call, triage and investigate live incidents with Datadog and MX investigation patterns, and take Incident Commander or supporting technical roles as the incident needs.
- After incidents, close the detection loop (gap packs, new monitors, dashboards) so the pager gets quieter over time.
- Run high-value launch and production-readiness reviews as a checkpoint, not a permanent staffing model.
Basic Requirements:
- BS in Computer Science or equivalent experience
- 5+ years running production observability, SRE, or DevOps. Datadog preferred; strong Grafana/Prometheus, Splunk, or New Relic experience counts if you can ramp on Datadog fast.
- Automation-first engineering in Python, Bash, Go, and/or Terraform, plus Kubernetes proficiency. You encode monitoring standards as code rather than clicking the UI.
- AI- and workflow-literate. You've used or built scripted and AI-assisted workflows to scale reviews, audits, and docs.
- Alerting and SLO strategy: burn-rate and error-budget thinking, with a track record of cutting alert fatigue on evidence.
- Distributed-systems debugging across microservices: latency, connection pools, queues, and cascading failure on Kubernetes and bare metal, with NATS, RabbitMQ, Postgres, and Redis.
- Shared on-call, Incident Commander-capable. You've run or supported incident bridges and written postmortems, and you'll take shifts on the shared IR & Observability rotation.
Preferred Requirements:
- Fintech experience with MX-like architectures
- Google SRE practices: toil elimination, incident management, automation for self-healing
- Cross-functional influence without authority. You've improved teams that don't report to you.
- Governance and reporting: you can produce a monthly health and compliance report leadership reads (orphans, stale entries, gaps, trends).
- OpenTelemetry instrumentation
- Datadog cost optimization at scale (cardinality, log indexing, sampling)
- Incident response platforms (incident.io, PagerDuty, OpsGenie); prior formal Incident Commander experience
- Golang and Ruby on Rails (the MX stack)
What Success Looks Like:
- By six months, you're a full participant on the shared on-call rotation and a capable Incident Commander on live SEVs, teams you've engaged have alerts and dashboards that answer "what's broken and where do I look?", and the monthly health report runs largely on its own. By twelve months, incidents get caught earlier because of instrumentation the loop added, new services get baseline observability from a self-serve template on day one, and no team depends on a shepherd for day-one coverage.