Senior/Staff System Research Engineer – LLM Inference Optimization
Bellevue, United States · Hybrid · Full-time
- Posted 2w ago
- From Snowflake’s careers page
- Location
- Bellevue, United States
- Work mode
- Hybrid
- Type
- Full-time
- Level
- Staff
- Experience
- 5+ years
- Department
- Engineering
Opens the listing on jobs.ashbyhq.com
Let the right jobs find you
In your inbox every Wednesday and SaturdayPersonalised suggestions from verified career pages, matched to your role, location, level and skills.
About the role
Responsibilities
-
Design and develop high-performance LLM inference systems, spanning distributed serving, runtime systems, GPU execution, and performance-critical kernels.
-
Develop novel techniques to improve inference latency, generation speed, throughput, memory efficiency, scalability, and cost.
-
Explore advanced inference techniques including speculative and parallel decoding, prefill/decode disaggregation, adaptive parallelism, continuous batching and scheduling, KV-cache management, quantization, and communication optimization.
-
Develop adaptive and intelligent inference systems that automatically optimize execution for new model architectures, hardware platforms, workload characteristics, and deployment environments.
-
Apply AI-driven and AI-native approaches to systems engineering, including automated profiling, bottleneck identification, configuration search, code generation, experimentation, runtime strategy selection, debugging, and performance tuning.
-
Independently identify high-impact performance and systems problems, formulate hypotheses, prototype solutions, and drive promising ideas from research through production.
-
Design distributed inference strategies across GPUs and nodes, including tensor, sequence, pipeline, data, and expert parallelism.
-
Develop efficient approaches for multi-model serving, dynamic resource management, model loading and swapping, and workload-aware scheduling.
-
Analyze and optimize GPU kernels and operators for attention, MoE, communication, and other performance-critical model components.
-
Explore model-system co-design, including model or post-training techniques that unlock substantially more efficient inference.
-
Profile and benchmark end-to-end workloads to identify bottlenecks across compute, memory, communication, networking, scheduling, and model execution.
-
Collaborate closely with model researchers, infrastructure teams, and product teams to deploy research innovations in production.
-
Open-source and publish innovations through technical blogs and top-tier systems and machine learning conferences.
Requirements
-
Bachelor’s degree in Computer Science, Electrical Engineering, or a related field. A Master’s degree or PhD is preferred.
-
5+ years of experience in one or more of the following areas: LLM inference systems, distributed AI systems, GPU systems, or high-performance computing.
-
Strong understanding of modern LLM inference architectures and the performance tradeoffs involved in serving large-scale models.
-
Hands-on experience with modern LLM inference and serving frameworks, such as vLLM, SGLang, TensorRT-LLM, or similar systems.
-
Experience designing, extending, or optimizing inference runtimes, including areas such as scheduling, batching, KV-cache management, distributed execution, parallelism, speculative decoding, or disaggregated serving.
-
Strong understanding of GPU architectures and experience with CUDA, Triton, or similar GPU programming environments.
-
Experience with performance-oriented libraries and frameworks such as CUTLASS, cuBLAS, cuDNN, or related technologies.
-
Experience profiling and diagnosing end-to-end system performance using Nsight Systems, Nsight Compute, or equivalent tools.
-
Demonstrated ability to operate as an independent problem identifier and solver—recognizing important problems with limited direction, defining the right technical questions, and driving solutions through ambiguity.
-
Strong ability to work across model, runtime, distributed system, and hardware layers and reason about end-to-end performance tradeoffs.
-
Experience using AI-native engineering approaches to accelerate software development, experimentation, debugging, optimization, or system adaptation is a strong plus.
-
Excellent communication skills and the ability to collaborate effectively across research, engineering, and product teams.
Snowflake is growing fast, and we’re scaling our team to help enable and accelerate our growth. We are looking for people who share our values, challenge ordinary thinking, and push the pace of innovation while building a future for themselves and Snowflake.
How do you want to make your impact?
For jobs located in the United States, please visit the job posting on the Snowflake Careers Site for salary and benefits information: careers.snowflake.com
Skills they ask for
Pick one to see other roles that ask for it.
About Snowflake
Cloud data, analytics and AI platformSnowflake provides a managed cloud platform for data engineering, analytics, AI, and building and sharing data applications.
See all 170 roles at SnowflakeMore roles at Snowflake
See all 170- Account EngineerChicagoEngineeringChicago, United States3h
- Account Executive, Enterprise AcquisitionUnited States · RemoteSales · RemoteUnited States9h
- Brand DesignerMenlo ParkMarketingMenlo Park, United States10h
- Director of Engineering - Traffic & NetworkingMenlo Park · Director · On-siteEngineering · Director · On-siteMenlo Park, United States1d
Let the right jobs find you
In your inbox every Wednesday and SaturdayPersonalised suggestions from verified career pages, matched to your role, location, level and skills.