Software Engineer, Compute Infrastructure
San Francisco, United States · Hybrid · Full-time
- Posted 1mo ago
- From OpenAI’s careers page
- Location
- San Francisco, United States
- Work mode
- Hybrid
- Type
- Full-time
- Department
- Software Development
Opens the listing on jobs.ashbyhq.com
Let the right jobs find you
In your inbox every Wednesday and SaturdayPersonalised suggestions from verified career pages, matched to your role, location, level and skills.
About the role
About the Team:
Compute Infrastructure builds the platform that turns enormous amounts of compute into a reliable engine for frontier AI. We design, provision, schedule, operate, and optimize the systems that connect accelerators, CPUs, networks, storage, data centers, orchestration software, agent infrastructure, developer tools, and observability into one coherent experience for researchers and product teams.
About the Role
We are looking for engineers who want to build the compute platform behind OpenAI's research and products. You may not be the strongest in low-level systems, high-performance computing, distributed infrastructure, reliability, CaaS, agent infrastructure, developer platforms, tooling, or the user experience around infrastructure. What matters is that you can reason carefully about complex systems, write durable software, and raise the quality and velocity of the people around you.
Where you might work
-
Compute Foundations: Build the low-level platform primitives that make heterogeneous hardware, providers, and data centers repeatable, automatable, and operable at scale.
-
Fleet / Orchestration: Turn raw capacity into reliable, efficient clusters and scheduling systems that researchers and product teams can use with minimal friction and great experience.
-
Core Network Engineering: Build and operate the high-performance networking fabrics, protocols, and observability needed for the largest training and serving workloads.
-
Hardware Health and Observability: Detect, diagnose, remediate, and prevent hardware and fleet-health issues so usable compute stays high across providers and accelerator generations.
-
Storage: Build scalable, performant, durable storage abstractions that keep data movement and storage access from becoming a bottleneck to research or products.
-
Agent Infrastructure: Build sandboxed execution infrastructure for agentic workloads across research and production, with strong isolation, reliability, and scale.
In this role, you will:
-
Build and deeply optimize reliable system software for large-scale compute systems that run some of the world's most demanding AI workloads
-
Design and operate infrastructure across accelerators, CPUs, NICs, switches, networking protocols, storage, data centers, cluster orchestration, scheduling, and fleet health
-
Profile, benchmark, and optimize training workloads across compute, memory, storage, networking, NCCL and collective communication, and cluster scheduling bottlenecks
-
Create hardware-aware automation that makes provisioning, firmware and driver upgrades, incident response, and day-to-day operations faster and less error-prone
-
Build CaaS, agent infrastructure, profiling, observability, benchmarking, and platform tools that help researchers, product engineers, and operators launch, debug, and optimize workloads with less friction
-
Turn operational lessons into better systems, stronger abstractions, and clearer ownership boundaries across teams
-
Collaborate across research, engineering, security, networking, hardware, and data center teams to make compute capacity more capable and easier to use
You might thrive in this role if you:
-
Have built or operated distributed systems, infrastructure platforms, high-performance computing environments, large-scale networking systems, Kubernetes clusters, developer tools, or production systems with demanding reliability requirements
-
Enjoy working across layers of the stack and are comfortable moving between software, hardware, networking, systems performance, reliability, and user needs
-
Care about making complex infrastructure understandable, observable, and usable for the people depending on it
-
Can diagnose hard problems under real operational pressure while still investing in long-term engineering quality
-
Like building leverage for others, whether through APIs, automation, debugging tools, CaaS and agent infrastructure primitives, workflow improvements, or better platform abstractions
-
Are motivated by scale, efficiency, reliability, and disciplined measurement through benchmarks, profiles, and production evidence
-
Communicate clearly, take ownership, and work well with teams whose constraints and goals differ from your own
Qualifications
-
Strong software engineering skills and experience building, operating, or improving production infrastructure systems
-
Experience in one or more relevant areas such as distributed systems, operating systems, networking protocols, RDMA, NCCL or collective communication, storage, Kubernetes, scheduling, observability, reliability engineering, high-performance computing, GPU infrastructure, CaaS, agent infrastructure, hardware-aware performance optimization, benchmarking, developer experience, or infrastructure tooling
-
Ability to debug complex system behavior across software, hardware, networking, and workload layers, then turn findings into robust improvements
-
Comfort with ambiguity, strong ownership, and a bias toward practical, durable solutions
-
Interest in working on infrastructure that directly enables frontier AI research and product impact
Skills they ask for
Pick one to see other roles that ask for it.
About OpenAI
AI research and deploymentOpenAI conducts AI research and develops products and platforms for consumers, developers and businesses.
See all 322 roles at OpenAIMore roles at OpenAI
See all 322- Market Research Lead, ChatGPTSan Francisco · Lead · HybridMarketing · Lead · HybridSan Francisco, United States8h
- Product Builder, SalesSan Francisco · HybridBusiness Operations · HybridSan Francisco, United States14h
- Account Director, CyberSan Francisco · HybridSales · HybridSan Francisco, United States23h
- Technical Accounting Lead, Ads RevenueSan Francisco · Lead · HybridFinance and Accounting · Lead · HybridSan Francisco, United States1d
Let the right jobs find you
In your inbox every Wednesday and SaturdayPersonalised suggestions from verified career pages, matched to your role, location, level and skills.