Data Center Hardware Quality & Reliability Engineer
San Francisco, United States · Hybrid · Full-time
- Posted 1w ago
- From OpenAI’s careers page
- Location
- San Francisco, United States
- Work mode
- Hybrid
- Type
- Full-time
- Level
- Senior
- Experience
- 8+ years
- Department
- Engineering
Opens the listing on jobs.ashbyhq.com
Let the right jobs find you
In your inbox every Wednesday and SaturdayPersonalised suggestions from verified career pages, matched to your role, location, level and skills.
About the role
About The Role
Own the end-to-end data-center hardware quality and reliability loop for OpenAI’s 3P infrastructure and 1P current and next-gen platforms. Turn field failures into quantified risk, fast containment, verified root cause, improved MQE/NPI and manufacturing-test coverage, accurate spares forecasts, and upstream changes that prevent recurrence.
The first hire must combine practical hardware/system understanding, reliability engineering, data fluency, and cross-functional technical leadership at data-center scale.
Key Responsibilities
- Build and govern the field-quality data model across telemetry, tickets, RMA/repair, FA, firmware, configuration, supplier, and manufacturing genealogy.
- Define AFR, ASR, DPPM, MTBF/MTTR, repeat-repair, NTF, repair-cycle-time, and forecast-versus-actual metrics with explicit denominators and uncertainty.
- Provide fleet-level macro views and unit/FRU/cohort-level micro views; detect shifts and bound affected populations.
- Lead systemic field-failure triage, containment, failure analysis, 8D/CAPA, risk assessment, corrective-action verification, and recurrence monitoring.
- Develop cohort, life-data, reliability-growth, and spare-demand projections by product, FRU, supplier, configuration, geography, and age.
- Partner with MQE and NPI to convert field mechanisms into manufacturing-test coverage, screening/stress profiles, diagnostics, control plans, DFR/DFS requirements, FMEA/FTA, mission profiles, FRU strategy, and qualification gates.
- Close the loop by verifying whether upstream changes reduce field recurrence.
- Define supplier/CM FA standards, field-data contracts, scorecards, escalation paths, and closure evidence.
- Provide serviceability, TCO, and spares inputs without owning inventory execution or procurement.
- Create concise executive decision packages: population at risk, exposure, confidence, options, cost/risk, and recommendation.
- Run the cross-functional reliability council and, as the team grows, mentor the 1P and 3P Field Quality Engineers.
Qualifications
- BS in electrical, mechanical, computer, materials, reliability engineering, physics, or equivalent experience; MS preferred.
- 8+ years in hardware quality/reliability, server/rack systems, or mission-critical infrastructure; 3+ years owning field-failure, RMA, or CAPA outcomes.
- Solid working understanding of hardware and system architecture across board, tray, rack, firmware, telemetry, manufacturing test, and fleet behavior; deep expertise in every subsystem is not required.
- Reliability statistics: censored life data, Weibull/Poisson/binomial methods, confidence bounds, MTBF/MTTR, and reliability growth.
- Hands-on FMEA/FTA, accelerated or reliability-demonstration testing, 8D/CAPA, FA, and corrective-action verification.
- Working proficiency with SQL and Python/R or equivalent analytics tools.
- Ability to influence design, validation, operations, suppliers/CMs, and senior leaders without direct authority.
Preferred Skills
- GPU/AI server platforms, liquid cooling, high-power delivery, high-speed networking, rack integration, or data-center operations.
- Design for serviceability: FRU boundaries, diagnostics, repair workflows, tooling/access, and spares policy.
- Qualification-to-field correlation and mission-profile development.
- ODM/CM/supplier experience: FA quality, audit, QBR, and corrective-action governance.
- Linux/BMC/IPMI/Redfish logs and fleet telemetry.
- Leadership of a cross-generation reliability program or launch-readiness gate.
Skills they ask for
Pick one to see other roles that ask for it.
About OpenAI
AI research and deploymentOpenAI conducts AI research and develops products and platforms for consumers, developers and businesses.
See all 328 roles at OpenAIMore roles at OpenAI
See all 328- Market Research Lead, ChatGPTSan Francisco · Lead · HybridMarketing · Lead · HybridSan Francisco, United States6h
- Product Builder, SalesSan Francisco · HybridBusiness Operations · HybridSan Francisco, United States13h
- Account Director, CyberSan Francisco · HybridSales · HybridSan Francisco, United States21h
- Technical Accounting Lead, Ads RevenueSan Francisco · Lead · HybridFinance and Accounting · Lead · HybridSan Francisco, United States1d
Let the right jobs find you
In your inbox every Wednesday and SaturdayPersonalised suggestions from verified career pages, matched to your role, location, level and skills.