Senior HPC Engineer, GPU Compute
Amsterdam · Remote · Full-time
- Posted 1mo ago
- From Nebius’s careers page
- Location
- Amsterdam
- Work mode
- Remote
- Type
- Full-time
- Level
- Senior
- Experience
- 5+ years
- Department
- Engineering
Apply on Nebius’s site
Opens the listing on careers.nebius.com
Let the right jobs find you
In your inbox every Wednesday and SaturdayPersonalised suggestions from verified career pages, matched to your role, location, level and skills.
About the role
The role:\n\nWe’re looking for a Senior HPC Cluster Engineer to join our team and play a key role in the development of our cutting-edge hyperscaler platform. The GPU & InfiniBand team is responsible for enhancing and optimizing the core components of our Cloud platform, with a specific focus on GPU computing, InfiniBand networks, and the KVM/QEMU stack. You’ll work closely with hardware virtualization and device emulation technologies, ensuring high performance and security in multi-GPU, HPC environments. The role involves analyzing, troubleshooting, and improving infrastructure to support new hardware, fine-tuning system performance, and automating fault detection and resolution in a complex system.\n\nIn this position, you will be responsible for:\n- Tuning the performance of GPU clusters and InfiniBand networks to ensure optimal operation in HPC and GPU-based environments.\n- Analyzing and troubleshooting the root cause of issues related to GPUs and InfiniBand networks, and proposing corrective actions.\n- Integrating new hardware into the existing infrastructure, including support for new GPU hardware through software stacks like Kubernetes, QEMU, and KVM.\n- Enhancing automation systems for proactive monitoring, detecting, and resolving issues in GPU and InfiniBand environments.\n- Configuring and managing GPU devices and InfiniBand fabrics, ensuring efficient and reliable operation.\n\nWe expect you to have:\n- 5+ years of professional experience in system-level software development (focused on performance optimization, low-level programming).\n- 3+ years of hands-on experience with Linux systems (administration, troubleshooting, and performance tuning).\n- In-depth understanding of server architecture, including PCIe devices, NICs, Linux OS/Kernel, and high-performance computing (HPC) systems.\n- Strong proficiency in one or more performance-oriented programming languages (C/C++, Go, Python).\n\nIt would be a plus if you have:\n- Experience with GPU end-to-end testing in a cluster environment using InfiniBand networking.\n- Proven track record of analyzing and optimizing the performance of HPC workloads (e.g., simulations, data analysis, AI/ML workloads).\n- Familiarity with RDMA, RoCE, and InfiniBand protocols for high-performance communication.\n- Background in Software-Defined Networking (SDN) and experience with HPC cluster networking.\n- Understanding of QEMU/KVM virtualization and managing virtualized environments.\n- Experience with deep learning frameworks such as PyTorch and TensorFlow, and their integration with HPC systems.\n- Familiarity with collective communication libraries like MPI and NCCL for distributed computing.\n\nWe conduct coding interviews as part of the process.\n\n#LI-LH2\n\n\n### Benefits & Perks:\n- Competitive compensation\n- Career growth and learning opportunities\n- Flexibility and ownership\n- Collaborative and innovative culture\n- Opportunity to work on impactful AI projects\n- International environment and talented teams\n\n### What's it like to work at Nebius:\nFast moving - Bold thinking - Constant growth - Meaningful impact - Trust and real ownership - Opportunity to shape the future of AI\n\n### Equal Opportunity Statement:\nNebius is an equal opportunity employer. We are committed to fostering an inclusive and diverse workplace and to providing equal employment opportunities in all aspects of employment. We do not discriminate on the basis of race, color, religion, sex (including pregnancy), national origin, ancestry, age, disability, genetic information, marital status, veteran status, sexual orientation, gender identity or expression, or any other characteristic protected by applicable law.\n\nApplicants must be authorized to work in the country in which they apply and will be required to provide proof of employment eligibility as a condition of hire. If you need accommodations during the application process, please let us know.
Skills they ask for
Pick one to see other roles that ask for it.
About Nebius
Cloud infrastructure for AINebius provides cloud infrastructure and services for building and scaling AI workloads.
See all 133 roles at NebiusMore roles at Nebius
See all 133- Principal ML Solutions Architect - Token FactoryUnited States · Principal · RemoteEngineering · Principal · RemoteUnited States12h
- Delivery Operations Manager - Token FactoryRemoteBusiness Operations · Remote12h
- Head of Strategic Partnerships, TavilyUnited States · Director · RemoteSales · Director · RemoteUnited States17h
- General Manager, Data Center (New Build)Independence · Director · On-siteInformation Technology · Director · On-siteIndependence, United States1d
Share this role
Let the right jobs find you
In your inbox every Wednesday and SaturdayPersonalised suggestions from verified career pages, matched to your role, location, level and skills.