Site Reliability Engineer

RapidAI

Full Time10+ yearsPosted 18 days ago

Let the right jobs find you

Get personalised suggestions from verified company career pages, matched to your role, location, level, and skills.

Overview

Position Type

Full Time

Experience

10+ years

Job Description

What You Do:

  • Own the availability, performance, and incident response for Rapid's production EKS clusters
  • Design and operate the full observability stack — metrics, logs, traces — with Open Telemetry as the foundation
  • Define and track SLOs/SLIs/error budgets; lead post-mortems and drive blameless culture
  • Build and maintain infrastructure-as-code using Terraform, Helm, and GitOps patterns
  • Partner with engineering to bake reliability in early — capacity planning, load testing, chaos engineering
  • Tune autoscaling, networking, and cost efficiency across AWS workloads
  • On-call rotation with the expectation you'll also fix the underlying cause, not just the alert

What We Looking For:

  • 10+ years in SRE, DevOps, or infrastructure engineering roles
  • Deep AWS expertise — EKS, EC2, VPC, IAM, RDS, S3, CloudWatch, and the surrounding ecosystem
  • Production Kubernetes experience at scale: multi-cluster, multi-tenant, real traffic
  • Hands-on Open Telemetry instrumentation and pipeline ownership (collectors, exporters, backends)
  • Strong foundation in Linux, networking, and distributed systems fundamentals
  • Experience with observability platforms (Prometheus, Grafana, Jaeger, or equivalents) Comfortable writing automation in Go, Python, or Bash — you reach for code when the GUI runs out
  • Startup mindset: you make decisions with incomplete information and iterate quickly

Required Skills

AwsEksAws Ec2Aws VpcAws IamAws RdsS3Cloud WatchKubernetesOpentelemetryTerraformHelmGitopsLinuxNetworking

About the Company

RapidAI

Bengaluru, India

Share This Job