DC Operations Manager - Critical Infrastructure
DC, United States · On-site · Full-time
- Posted 1mo ago
- From Nebius’s careers page
- Location
- DC, United States
- Work mode
- On-site
- Type
- Full-time
- Level
- Senior
- Experience
- 7+ years
- Department
- Engineering
Apply on Nebius’s site
Opens the listing on careers.nebius.com
Let the right jobs find you
In your inbox every Wednesday and SaturdayPersonalised suggestions from verified career pages, matched to your role, location, level and skills.
About the role
Roles and Responsibilities:
Critical Facilities Operations:
- Own day to day operations of electrical and mechanical systems including UPS, generators, switchgear, PDUs, chillers, CRAH/CRAC units, and cooling infrastructure
- Ensure high availability and uptime of all critical systems supporting data center operations
- Lead incident response for facility related events, including root cause analysis and corrective actions
- Monitor system performance via BMS/DCIM and drive improvements in reliability and efficiency
Maintenance & Vendor Management:
- Oversee preventive and corrective maintenance programs for all facility systems
- Manage and hold accountable third party vendors and service providers
- Ensure all maintenance activities follow SOPs, EOPs, and MOPs
- Drive standardization and continuous improvement of operational procedures
Capacity, Efficiency & Optimization:
- Partner with engineering and operations teams on capacity planning and infrastructure scaling
- Monitor and improve PUE and overall energy efficiency
- Identify and implement reliability and sustainability improvements across facilities
- Support high density environments, including air and liquid cooling strategies aligned with modern AI workloads
Compliance, Safety & Risk Management:
- Ensure compliance with local regulations, safety standards, and industry best practices
- Maintain strong adherence to HSE (Health, Safety, Environmental) standards
- Lead audits, inspections, and documentation for operational readiness
- Act as the primary point of contact for facility related regulatory interactions
Commissioning & Site Readiness (Light but Important):
- Support commissioning, testing, and handover of new or expanded infrastructure
- Participate in FAT/SAT and system validation for critical equipment
- Ensure smooth transition from construction to steady state operations
- Validate that systems are operationally ready with proper documentation and procedures
Cross-Functional Collaboration:
- Partner closely with:
- Data Center Operations teams
- Infrastructure & deployment engineers
- Network and hardware teams
- Act as the bridge between facilities and IT infrastructure, ensuring both layers operate seamlessly (a key distinction in Nebius environments)
- Support broader operational goals around scalability, reliability, and standardization
Requirements:
- 7–10+ years of experience in data center or mission critical facilities operations
- Strong hands on knowledge of critical infrastructure systems:
- Electrical distribution (UPS, generators, switchgear)
- Cooling systems (chillers, HVAC, CRAH/CRAC, liquid cooling exposure preferred)
- Experience operating in high availability environments with strict uptime requirements
- Proven experience managing vendors, maintenance programs, and incident response
- Familiarity with BMS/DCIM systems and operational monitoring
- Experience developing and enforcing SOPs, EOPs, and MOPs
- Strong understanding of safety, compliance, and regulatory requirements
Preferred:
- Experience in hyperscale, colocation, or AI/GPU driven data centers
- Exposure to high-density environments and advanced cooling solutions
- Experience supporting new site builds, commissioning, or expansions
- Basic understanding of IT infrastructure (rack, power distribution, hardware) to effectively partner with engineering teams
Key employee benefits:
- Health insurance: 100% company-paid medical, dental and vision coverage for employees and families.
- 401(k) plan: Up to 4% company match with immediate vesting.
- Parental leave: 20 weeks paid for primary caregivers, 12 weeks for secondary caregivers.
- Remote work reimbursement: Up to $85/month for mobile and internet.
- Disability & life insurance: Company-paid short-term, long-term and life insurance coverage.
Compensation
We offer competitive salaries, ranging from $190,000-$230,000 based on your experience
Skills they ask for
Pick one to see other roles that ask for it.
About Nebius
Cloud infrastructure for AINebius provides cloud infrastructure and services for building and scaling AI workloads.
See all 133 roles at NebiusMore roles at Nebius
See all 133- Principal ML Solutions Architect - Token FactoryUnited States · Principal · RemoteEngineering · Principal · RemoteUnited States11h
- Delivery Operations Manager - Token FactoryRemoteBusiness Operations · Remote11h
- Head of Strategic Partnerships, TavilyUnited States · Director · RemoteSales · Director · RemoteUnited States16h
- General Manager, Data Center (New Build)Independence · Director · On-siteInformation Technology · Director · On-siteIndependence, United States1d
Share this role
Let the right jobs find you
In your inbox every Wednesday and SaturdayPersonalised suggestions from verified career pages, matched to your role, location, level and skills.