Intermediate/ Senior Software Engineer - Cortex LLM Training Platform
Bellevue, United States · Hybrid · Full-time
- Posted 3w ago
- From Snowflake’s careers page
- Location
- Bellevue, United States
- Work mode
- Hybrid
- Type
- Full-time
- Level
- Senior
- Experience
- 6+ years
- Department
- Engineering
Opens the listing on jobs.ashbyhq.com
Let the right jobs find you
In your inbox every Wednesday and SaturdayPersonalised suggestions from verified career pages, matched to your role, location, level and skills.
About the role
Senior Software Engineer — Cortex Training
The Snowflake ML Platform team's mission is to let customers run their most demanding ML/AI workloads inside Snowflake. Cortex Training is our LLM post-training platform: it turns scarce, expensive GPU capacity into a simple, composable service, so customers can adapt open-weight foundation models to their own business problems while we handle the hard distributed-systems parts, including scheduling, orchestration, multi-node training and inference, fault tolerance, and throughput.
The platform already runs post-training at scale. Under the hood, it decouples GPU computation from the training loop and exposes it as primitive APIs that compose into everything from SFT to full RL workflows. You'll work alongside a team that ships fast & sweats reliability and the researchers behind DeepSpeed. We're looking for an engineer who thrives in the ML infrastructure layer and brings a solid understanding of LLMs and post-training to help us scale and grow it.
YOU WILL:
-
Design and build across the full stack — from the public training APIs and SDK through the control plane to the GPU data plane.
-
Scale the distributed systems that make GPU compute serverless — multi-tenant scheduling, placement, and capacity-aware routing across regional GPU pools, with fault tolerance built in.
-
Drive end-to-end performance at scale — keep the training, inference, and RL loops fast and the data plane responsive under heavy concurrent load, with GPUs kept saturated.
-
Productionize research building blocks — partner with Snowflake Research to turn state-of-the-art training and inference techniques into reliable, composable components customers can run at enterprise scale.
QUALIFICATIONS:
-
3 + years (Intermediate) | 6+ years (Senior) building and shipping production ML systems
-
Strong distributed systems and infrastructure foundation — designing scalable, fault-tolerant services and operating them on Kubernetes in production.
-
Familiarity with GPU and LLM infrastructure — e.g., PyTorch, DeepSpeed/FSDP, Ray, CUDA/NCCL, vLLM; able to debug across the data, infrastructure, and GPU layers.
-
Demonstrated ability to harden complex systems for reliability, throughput, and cost efficiency.
-
BS in Computer Science or a related field (MS/PhD a plus).
-
(Bonus) Hands-on LLM post-training / modeling experience — the strongest candidates pair deep infra skills with real post-training intuition.
Skills they ask for
Pick one to see other roles that ask for it.
About Snowflake
Cloud data, analytics and AI platformSnowflake provides a managed cloud platform for data engineering, analytics, AI, and building and sharing data applications.
See all 169 roles at SnowflakeMore roles at Snowflake
See all 169- Account Executive, Enterprise AcquisitionUnited States · RemoteSales · RemoteUnited States8h
- Brand DesignerMenlo ParkMarketingMenlo Park, United States9h
- Director of Engineering - Traffic & NetworkingMenlo Park · Director · On-siteEngineering · Director · On-siteMenlo Park, United States1d
- Sales Ops Manager, Partners & Specialists PlanningMenlo Park · HybridBusiness Operations · HybridMenlo Park, United States1d
Let the right jobs find you
In your inbox every Wednesday and SaturdayPersonalised suggestions from verified career pages, matched to your role, location, level and skills.