Site Reliability Engineer
RapidAI
Full Time10+ yearsPosted 18 days ago
Let the right jobs find you
Get personalised suggestions from verified company career pages, matched to your role, location, level, and skills.
Overview
Position Type
Full Time
Experience
10+ years
Job Description
What You Do:
- Own the availability, performance, and incident response for Rapid's production EKS clusters
- Design and operate the full observability stack — metrics, logs, traces — with Open Telemetry as the foundation
- Define and track SLOs/SLIs/error budgets; lead post-mortems and drive blameless culture
- Build and maintain infrastructure-as-code using Terraform, Helm, and GitOps patterns
- Partner with engineering to bake reliability in early — capacity planning, load testing, chaos engineering
- Tune autoscaling, networking, and cost efficiency across AWS workloads
- On-call rotation with the expectation you'll also fix the underlying cause, not just the alert
What We Looking For:
- 10+ years in SRE, DevOps, or infrastructure engineering roles
- Deep AWS expertise — EKS, EC2, VPC, IAM, RDS, S3, CloudWatch, and the surrounding ecosystem
- Production Kubernetes experience at scale: multi-cluster, multi-tenant, real traffic
- Hands-on Open Telemetry instrumentation and pipeline ownership (collectors, exporters, backends)
- Strong foundation in Linux, networking, and distributed systems fundamentals
- Experience with observability platforms (Prometheus, Grafana, Jaeger, or equivalents) Comfortable writing automation in Go, Python, or Bash — you reach for code when the GUI runs out
- Startup mindset: you make decisions with incomplete information and iterate quickly