Senior Software Engineer - Core Cloud Platform
San Francisco, United States · Hybrid · Full-time
- Posted 2w ago
- From Lambda’s careers page
- Location
- San Francisco, United States
- Work mode
- Hybrid
- Type
- Full-time
- Level
- Senior
- Experience
- 7+ years
- Department
- Software Development
Opens the listing on jobs.ashbyhq.com
Let the right jobs find you
In your inbox every Wednesday and SaturdayPersonalised suggestions from verified career pages, matched to your role, location, level and skills.
About the role
About the Role
As a Senior or Staff Software Engineer in Lambda’s Cloud Services Engineering organization, you will build and operate the distributed systems that power Lambda’s GPU cloud. Our teams own platform capabilities across compute control planes, managed Kubernetes, cloud APIs, identity and access, usage metering and billing, capacity and orchestration, reliability, and developer-facing infrastructure.
You will turn large-scale GPU infrastructure into reliable, secure, customer-facing cloud services by building APIs, workflows, stateful controllers, schedulers, and operational tooling. You will be full cycle engineer, owning systems through design, deployment, on-call, incident follow-through, and continuous improvement.
This role is a strong fit for engineers who enjoy cloud infrastructure, distributed systems, operational excellence, and solving ambiguous problems across software and infrastructure boundaries. Senior engineers lead complex work within a team or domain; Staff engineers additionally shape cross-team architecture and make other teams more effective.
What You’ll Do
-
Design, build, and operate services, APIs, control planes, and platform capabilities that power Lambda’s AI cloud.
-
Solve distributed-systems problems involving state, consistency, concurrency, scheduling, failure recovery, and safe lifecycle management.
-
Own the full engineering lifecycle: problem framing, architecture, implementation, testing, rollout, observability, on-call, and continuous improvement.
-
Improve system availability, latency, throughput, efficiency, security, and operability as Lambda grows by orders of magnitude.
-
Turn incidents and near misses into durable engineering improvements, including better automation, testing, guardrails, and backstops.
-
Work across product, infrastructure, networking, storage, security, and SRE teams to resolve dependencies and deliver the right outcome for customers.
-
Use AI-assisted development tools with judgment: accelerate exploration and implementation while independently verifying correctness, security, and maintainability.
-
Contribute to technical standards, design and code reviews, and mentorship; at Staff level, lead cross-team architecture and raise the technical ceiling of the organization.
Skills they ask for
Pick one to see other roles that ask for it.
About Lambda
AI compute in the cloudLambda provides cloud GPU compute, clusters and AI infrastructure for researchers, startups and enterprises.
See all 62 roles at LambdaMore roles at Lambda
See all 62- Network Planning & Design ManagerSan Francisco · HybridEngineering · HybridSan Francisco, United States5h
- Network Infrastructure Delivery ManagerSan Francisco · HybridBusiness Operations · HybridSan Francisco, United States5h
- Network AnalystSan Francisco · Entry Level · HybridBusiness Operations · Entry Level · HybridSan Francisco, United States5h
- Staff Product Manager - ComputeBellevue · Staff · HybridProduct Management · Staff · HybridBellevue, United States8h
Let the right jobs find you
In your inbox every Wednesday and SaturdayPersonalised suggestions from verified career pages, matched to your role, location, level and skills.