HPC Systems Architect
San Jose, United States · Hybrid · Full-time
- Posted 3w ago
- From Lambda’s careers page
- Location
- San Jose, United States
- Work mode
- Hybrid
- Type
- Full-time
- Level
- Staff
- Experience
- 7+ years
- Department
- Data and Analytics
Apply on Lambda’s site
Opens the listing on jobs.ashbyhq.com
Let the right jobs find you
In your inbox every Wednesday and SaturdayPersonalised suggestions from verified career pages, matched to your role, location, level and skills.
About the role
What You’ll Do
- Architect and define scalable compute platforms optimized for AI/ML, simulation, and high-throughput workloads.
- Develop compute system standards and design patterns to ensure consistency, performance, and maintainability across infrastructure.
- Evaluate emerging CPU, GPU, and accelerator technologies, owning architectural tradeoff decisions that impact compute density, power, cooling, and total cost.
- Collaborate with product and engineering teams to map workload requirements to compute platform capabilities across bare metal and cloud deployments.
- Experience converting ambiguous business or customer needs into measurable platform requirements, technical specifications, acceptance criteria, and architecture decisions.
- Define compute platform roadmaps and architectural reference designs that guide hardware selection, firmware baselines, rack-level, and cluster design.
- Act as a technical lead during new platform introductions, guiding validation and performance characterization efforts.
- Mentor systems engineers and cross-functional stakeholders on compute performance tuning, sizing, and architectural decisions.
You
- Proven experience (7+ years) architecting large-scale 10k-100k+ GPU HPC or cloud compute platforms.
- Deep knowledge of CPU/GPU architectures, memory hierarchies, and accelerator topologies.
- Experience designing systems around high-bandwidth, low-latency fabrics (NVLink, InfiniBand, and RoCE).
- Strong understanding of system performance tuning, resource scheduling, thermal and power optimization, and compute lifecycle management.
- Comfortable working across hardware and software boundaries, especially at the intersection of compute architecture, OS behavior, and orchestration layers.
- Skilled at balancing architectural tradeoffs for density, power efficiency, cooling, and performance.
- Strong analytical and communication skills, with a track record of influencing technical strategy across teams.
- Strong ownership and can do attitude, self-starter who feels comfortable working in ambiguity.
Nice to Have
- Hands-on experience with AI/ML workloads and their compute performance characteristics.
- Familiarity with orchestration tools used in HPC. (Slurm, Kubernetes, etc)
- Experience with virtualization technologies, specifically GPU virtualization.
- Exposure to hardware validation, vendor collaboration, and long-term OEM roadmap alignment.
- Background in compute telemetry, real-time performance profiling, or large-scale A/B infrastructure testing.
Skills they ask for
Pick one to see other roles that ask for it.
- AI ml workloads
- Gpu architecture
- Memory hierarchy
- Accelerator topologies
- High bandwidth low latency fabrics
- Performance tuning
- Resource scheduling
- Thermal and power optimization
- Lifecycle management
- Orchestration tools
- Virtualization
- Hardware validation
- Compute telemetry
- Real time performance profiling
- Large scale infrastructure
About Lambda
AI compute in the cloudLambda provides cloud GPU compute, clusters and AI infrastructure for researchers, startups and enterprises.
See all 62 roles at LambdaMore roles at Lambda
See all 62- Network Planning & Design ManagerSan Francisco · HybridEngineering · HybridSan Francisco, United States5h
- Network Infrastructure Delivery ManagerSan Francisco · HybridBusiness Operations · HybridSan Francisco, United States5h
- Network AnalystSan Francisco · Entry Level · HybridBusiness Operations · Entry Level · HybridSan Francisco, United States5h
- Staff Product Manager - ComputeBellevue · Staff · HybridProduct Management · Staff · HybridBellevue, United States8h
Share this role
Let the right jobs find you
In your inbox every Wednesday and SaturdayPersonalised suggestions from verified career pages, matched to your role, location, level and skills.