Performance Engineer, Inference Engine
San Francisco, United States · On-site
- Posted 1mo ago
- From Anthropic’s careers page
- Location
- San Francisco, United States
- Work mode
- On-site
- Department
- Engineering
Opens the listing on job-boards.greenhouse.io
Let the right jobs find you
In your inbox every Wednesday and SaturdayPersonalised suggestions from verified career pages, matched to your role, location, level and skills.
About the role
About the Role
Anthropic's inference engine is the software between the accelerator kernels and the routing layer. It manages the entire token path in between: batching requests, laying the model out across chips, managing memory for weights and activations, coordinating every forward pass, and managing model state across requests. Built in-house, it runs on all of our accelerator platforms, serving Claude to millions of users and running our research workloads.
You will work on building and optimizing this system at Anthropic scale: improving throughput, cost, reliability, and latency across all accelerator and cloud platforms. You are intimately familiar with the hardware and bandwidth numbers (FLOPs, HBM, PCIe, RDMA, network links, etc.) and can model a problem quickly: where the time and bytes go, and what sets the bound. The role is deeply technical and high-impact, and suits engineers who enjoy working across accelerator programming, high-performance systems that seamlessly coordinate between host and device, and large-scale distributed systems. Familiarity with the transformer architecture is a plus.
Minimum Qualifications
- A working mental model of LLM inference: how prefill and decode land on an accelerator's compute, memory, and interconnect, and what the host is doing meanwhile
- Proven quick learner: ramped fast in deep, unfamiliar systems and shipped consequential changes quickly
- Strong systems programming (Rust, C++, or similar), with care for code quality and tests
- Analytical about performance: observe and profile first, form a hypothesis, test it, then change the code and measure again
- Low ego: ask the naive question, take feedback well, pick up slack outside your job description
- Enjoy pair programming (we love to pair!) and care about the societal impacts of your work
Preferred Qualifications
- Experience inside an LLM serving engine and a sense of where its abstractions strain
- GPU/Accelerator programming
- OS internals
- Language modeling with transformers
- Experience building an allocator, cache, scheduler, or high-bandwidth transport
- Fluency in Rust
- Experience making systems reproducible: determinism, replay, property-based tests
Skills they ask for
Pick one to see other roles that ask for it.
About Anthropic
AI research and safetyAnthropic researches and builds AI systems, with work spanning model development, safety, and societal impacts.
See all 312 roles at AnthropicMore roles at Anthropic
See all 312- Head of Programmatic Customer SuccessSan Francisco · Director · On-siteCustomer Service · Director · On-siteSan Francisco, United States4h
- Cybersecurity Sales SpecialistSan Francisco · On-siteSales · On-siteSan Francisco, United States5h
- Podcast Media ManagerSan Francisco · On-siteMarketing · On-siteSan Francisco, United States7h
- Data Engineer, ProductSan Francisco · On-siteData and Analytics · On-siteSan Francisco, United States8h
Let the right jobs find you
In your inbox every Wednesday and SaturdayPersonalised suggestions from verified career pages, matched to your role, location, level and skills.