Site Reliability Engineer
TestMu AI (LambdaTest)
Full Time1+ yearsPosted 25 days ago
Let the right jobs find you
Get personalised suggestions from verified company career pages, matched to your role, location, level, and skills.
Overview
Position Type
Full Time
Experience
1+ years
Job Description
The Role What You'll Do:
- This is a genuine DevOps/SRE role — not config management or ticket ops. You'll write code, own incidents end-to-end, and build custom automation solutions on live production systems.
- Cloud Operations: Own and scale services on AWS — EKS, EC2, SQS, ECR, Route 53, ALB/NLB.
- Kubernetes & Containers: Manage production workloads on EKS; implement Karpenter and KEDA for event-driven and node autoscaling.
- CI/CD: Build and optimize pipelines using Jenkins, Docker Buildx, ECR caching, ArgoCD, and Helm.
- Observability & Incident Management: Own APM, monitoring, and production debugging using New Relic, Sumo Logic, Prometheus, or Grafana.
- Automation & Scripting: Write custom code for logical and automation solutions — not just run commands on machines.
- IaC: Use Terraform for provisioning; understand state management and concurrent apply risks in team environments.
You'll Thrive Here If You Have | Must-Haves:
- Experience: 1–3 years of hands-on DevOps or Cloud Infrastructure experience.
- Cloud: Production experience on AWS — EKS, SQS, ECR, Route 53, ALB/NLB.
- Containers: Docker and Kubernetes in production environments.
- Observability: Hands-on with at least one APM tool — New Relic, Sumo Logic, Prometheus, or Grafana.
- Incident Management: Real production incident experience — must articulate root cause, not just symptoms.
- Coding: Genuine programming ability in any language — this role builds custom solutions, not just runs CLI commands.
- Mindset: Bridges dev and DevOps thinking; can explain why architectural choices were made, not just what was done.
Good to Have:
- Golang or Java — backend coding ability is a strong plus.
- KEDA and Karpenter — event-driven and node autoscaling experience.
- ArgoCD, Helm, or Istio — GitOps and service mesh exposure.
- Kafka or SQS — messaging and queue configuration knowledge.
- System design awareness at the services layer — how services interact, not just surface-level ops.
What We Offer:
Direct exposure to production systems at scale, ownership from day one, and a clear growth path in Platform and SRE Engineering.
Apply now to build infrastructure that scales.