Senior Platform Engineer - Core Infrastructure
San Francisco, United States · Hybrid · Full-time
- Posted 1mo ago
- From Lambda’s careers page
- Location
- San Francisco, United States
- Work mode
- Hybrid
- Type
- Full-time
- Level
- Senior
- Experience
- 5+ years
- Department
- Engineering
Opens the listing on jobs.ashbyhq.com
Let the right jobs find you
In your inbox every Wednesday and SaturdayPersonalised suggestions from verified career pages, matched to your role, location, level and skills.
About the role
What You’ll Do
-
Architect, deploy, and operate Kubernetes clusters across Lambda's bare-metal datacenters.
-
Build and maintain automation for cluster lifecycle management: provisioning, upgrades, and scaling.
- Similar to SRE-level coding abilities (small scripts, k8s operators/CRDs a plus)
-
Own the reliability, performance, and security of Kubernetes workloads in production.
-
Implement observability, logging, and alerting for clusters and critical workloads.
-
Partner with product teams to design scalable, cloud-native services.
-
Set the standards for resource management, networking, and RBAC across the platform.
-
Lead incident response, root-cause analysis, and post-mortems for platform issues.
-
Mentor junior engineers and raise the bar for platform engineering across the org.
You
-
5+ years in Platform, Infrastructure, or SRE roles, including running Kubernetes in production at scale.
-
Deep knowledge of Kubernetes internals and day-2 operations (upgrades, scaling, troubleshooting).
-
Strong with Helm, Kustomize, or similar, and GitOps-based delivery.
-
Proficient with Linux systems and understanding of system-level operations.
-
Proficient with infrastructure-as-code (Terraform, Pulumi, or equivalent).
-
Solid grounding in networking, service meshes, and container runtimes.
-
Hands-on with observability stacks (Prometheus, Grafana, OpenTelemetry).
-
Strong coding skills in Go or Python for automation and tooling.
-
Practical security experience: network policies, secrets management, and image scanning.
Nice to Have
-
Experience with multi-cluster, multi-cloud, or hybrid environments.
-
Knowledge of GPU scheduling, HPC workloads, or ML/AI infrastructure.
-
Experience with workflow orchestration / durable execution frameworks (Temporal, Cadence, or Argo Workflows).
-
Exposure to cost optimization and capacity planning for large clusters.
-
Contributions to CNCF or Kubernetes open-source projects.
-
CKA/CKS certification.
Skills they ask for
Pick one to see other roles that ask for it.
About Lambda
AI compute in the cloudLambda provides cloud GPU compute, clusters and AI infrastructure for researchers, startups and enterprises.
See all 64 roles at LambdaMore roles at Lambda
See all 64- Network Planning & Design ManagerSan Francisco · HybridEngineering · HybridSan Francisco, United States4h
- Network Infrastructure Delivery ManagerSan Francisco · HybridBusiness Operations · HybridSan Francisco, United States4h
- Network AnalystSan Francisco · Entry Level · HybridBusiness Operations · Entry Level · HybridSan Francisco, United States4h
- Staff Product Manager - ComputeBellevue · Staff · HybridProduct Management · Staff · HybridBellevue, United States6h
Let the right jobs find you
In your inbox every Wednesday and SaturdayPersonalised suggestions from verified career pages, matched to your role, location, level and skills.