Capacity Operations Manager
San Francisco, United States · Hybrid · Full-time
- Posted 1mo ago
- From Baseten’s careers page
- Location
- San Francisco, United States
- Work mode
- Hybrid
- Type
- Full-time
- Level
- Senior
- Experience
- 5+ years
- Department
- Business Operations
Opens the listing on jobs.ashbyhq.com
Let the right jobs find you
In your inbox every Wednesday and SaturdayPersonalised suggestions from verified career pages, matched to your role, location, level and skills.
About the role
The Role
We're looking for a hands-on Operations Manager to own the operational and analytical supply side of our GPU fleet. Key focus areas: GPU fleet lifecycle, health, observability, utilization monitoring, and remediation across our neocloud and bare metal environments.
We contract for a fixed amount of compute capacity. GPUs drift from healthy to unhealthy over time, and this role minimizes that downtime to keep the maximum number of GPUs healthy at any given moment.
This is an operator role, not people management. You'll drive execution through clear processes, metrics, reporting, vendor coordination, and cross-functional alignment.
Responsibilities
- Drive suppliers to keep the maximum amount of the GPU fleet online and healthy.
- Maintain a live reconciliation of contracted vs. provisioned vs. healthy vs. utilized capacity, broken out by supplier and by cluster maximizing the number of healthy GPUs.
- Supplier-attributed fleet health accountability: own replacement SLAs, mean time to repair (MTTR), and RMA cycle times for every in-scope supplier.
- SLA monitoring, credit claims, and remedy enforcement: track SLA performance against contract terms, file and pursue credit claims, and drive remediation plans when suppliers fall short.
- Drive internal communications where suppliers need to perform maintenance to ensure all Baseten stakeholders are aware of activities that impact availability.
Requirements
- 5 to 10+ years within infrastructure working within the compute lifecycle to maximize functional compute, ideally in a hyperscale, cloud, or large-scale compute environment.
- Direct experience managing GPU, server, or data center hardware supplier relationships. You understand fleet health, RMA processes, and how contracted capacity differs from delivered capacity.
- Highly analytical. You should be comfortable pulling your own data, building your own reports, and generating insights without waiting on someone else to hand you a dashboard.
- Comfortable with ambiguity. Part of the job is figuring out what should exist and building it.
- Strong cross-functional collaboration skills. You'll work closely with finance, infrastructure/engineering, legal, and security on a regular basis.
Preferred Qualifications:
- Experience at a hyperscaler, neo cloud provider, or AI infrastructure company
- Familiarity with GPU hardware lifecycles (NVIDIA H100/H200/GB200 class systems), power/thermal constraints, and supply chain dynamics for compute.
- Experience running formal supplier corrective actions.
Skills they ask for
Pick one to see other roles that ask for it.
About Baseten
Deploy and run AI models in productionBaseten provides infrastructure for deploying and serving AI models in production.
See all 54 roles at BasetenMore roles at Baseten
See all 54- GRC ManagerSan Francisco · HybridLegal and Compliance · HybridSan Francisco, United States2d
- Post-Training Applied ResearcherSan Francisco · HybridResearch and Development (R&D) · HybridSan Francisco, United States2d
- Engineering Manager - Inference PerformanceSan Francisco · Senior · HybridEngineering · Senior · HybridSan Francisco, United States3d
- Revenue AnalystSan Francisco · HybridData and Analytics · HybridSan Francisco, United States3d
Let the right jobs find you
In your inbox every Wednesday and SaturdayPersonalised suggestions from verified career pages, matched to your role, location, level and skills.