HPC Infrastructure Engineer

ElevenLabs

Full TimeNot specifiedPosted 10 days ago

Let the right jobs find you

Get personalised suggestions from verified company career pages, matched to your role, location, level, and skills.

Overview

Position Type

Full Time

Experience

Not specified

Job Description

About the role

Every model we train runs on infrastructure this role owns. We operate NVIDIA GPU clusters across bare metal and rented capacity, and we're looking for an engineer to join our small research infrastructure team and make that compute fast, reliable, and boring - in the best sense. When the clusters just work, research moves faster. Your impact is measured directly in training throughput and researcher velocity.

What you’ll be doing

  • Operate and improve our GPU fleet end to end: provisioning, scheduling, monitoring, upgrades, capacity planning
  • Build automation that keeps the fleet healthy without human intervention — node health checks, automated draining and remediation, burn-in pipelines for new capacity
  • Own the stack beneath the training code: OS images, NVIDIA drivers, CUDA, container runtimes, NCCL, high-speed networking (InfiniBand/RoCE)
  • Run and tune job scheduling (Slurm or similar) so researchers get compute fairly and fast
  • Build and maintain high-performance storage for datasets and checkpoints
  • Hunt down performance problems: stragglers, degraded links, thermal issues, flaky GPUs — and fix the class of problem, not just the instance
  • Evaluate rented GPU capacity: benchmark it, validate it, hold providers to their SLAs
  • Hands-on hardware work when it’s needed: racking, cabling, diagnostics, coordinating with datacenter staff and vendors
  • Keep clusters secure by default: access control, network isolation, secrets

Requirements

  • Have run large-scale Linux server or GPU environments in production and enjoy both building and operating
  • Know the NVIDIA stack well — drivers, CUDA, NCCL, DCGM — or have deep systems experience and learn hardware stacks fast
  • Are comfortable with bare-metal environments, server hardware, and high-speed networking
  • Write solid automation in Python and/or Bash, with IaC tools like Ansible or Terraform
  • Are happy digging into noisy data (metrics, logs, PromQL) to find what’s actually wrong
  • Like owning real scope end to end and being the person others rely on
  • Don’t consider any task above or beneath you — datacenter trips included

Nice to have

  • Experience supporting ML training workloads from the infra side (distributed training failure modes, checkpointing patterns)
  • Experience evaluating and working with GPU cloud providers
  • Parallel filesystems (WEKA, VAST, etc) or large-scale object storage
  • BMC/IPMI/Redfish automation, PXE provisioning at scale
  • Power and cooling awareness for dense GPU deployments

How we work

Small team, high trust, minimal bureaucracy. We automate aggressively so on-call is sane, and we fix root causes so the same page never fires twice. You’ll work directly with the researchers whose jobs run on your clusters — short feedback loops, no ticket queues between you and impact.

Required Skills

LinuxGpuNvidiaCudaNcclDcgmPythonBashAnsibleTerraformProm QlSlurm

About the Company

ElevenLabs

United States

Share This Job