Senior ML Solutions Architect
Remote · Full-time
- Posted 1mo ago
- From Nebius’s careers page
- Work mode
- Remote
- Type
- Full-time
- Level
- Senior
- Experience
- 5+ years
- Department
- Engineering
Opens the listing on careers.nebius.com
Let the right jobs find you
In your inbox every Wednesday and SaturdayPersonalised suggestions from verified career pages, matched to your role, location, level and skills.
About the role
This position sits within Nebius Token Factory, our serverless platform for running and customizing open-source LLMs in production. Token Factory allows for serverless inference and fine-tuning (LoRA, full FT, RFT) backed by in-house optimizations like custom speculative decoding, quantization, cache-aware routing and dedicated endpoints. Customers come to us to move from prototype to scaled production without the cost and complexity of building and tuning their own inference stack. Responsibilities include: Optimize LLM inference across various modalities to drive business value and support customer goals. Provide support in supervised and reinforcement learning fine-tuning to maximize model quality for the customers. Design and implement LLM-based solutions using Nebius Token Factory’s inference services. Build production-ready applications leveraging our serverless LLM APIs, including multimodal models (text, vision, audio) and domain-specific models. Provide technical expertise in prompt engineering, RAG architectures and model selection. Collaborate with product and engineering teams to surface customer feedback and shape the platform roadmap. Guide customers in scaling from POC to production with a focus on performance, reliability, and cost efficiency. Qualifications include: 5+ years of experience in ML/AI systems, with at least 2 years focused on LLMs and generative AI; deep knowledge of the LLM ecosystem, model architectures and fine-tuning approaches; hands-on experience with running LLMs in production and LLM fine-tuning (SFT/LoRA) and data preparation/curation; RL-based fine-tuning is a strong plus; LLM evaluation: building benchmarks and eval pipelines; inference frameworks and libraries (e.g., vLLM, SGLang, TensorRT-LLM, Transformers); deploying LLM-powered applications using APIs from OpenAI, Anthropic, or open-source models; strong Python programming; excellent communication skills. Preferred: experience with inference frameworks, multimodal models, DevOps tools (Docker, Kubernetes), and open-source contributions. Technical stack: Python; vLLM, TensorRT-LLM, SGLang, Transformers; Kubernetes, Docker, Git; AWS SageMaker/Bedrock, GCP Vertex AI, Azure ML. Benefits include competitive compensation, career growth, flexibility, collaborative culture, and opportunity to work on impactful AI projects. Equal Opportunity Nebius is an equal opportunity employer.
Skills they ask for
Pick one to see other roles that ask for it.
About Nebius
Cloud infrastructure for AINebius provides cloud infrastructure and services for building and scaling AI workloads.
See all 133 roles at NebiusMore roles at Nebius
See all 133- Principal ML Solutions Architect - Token FactoryUnited States · Principal · RemoteEngineering · Principal · RemoteUnited States10h
- Delivery Operations Manager - Token FactoryRemoteBusiness Operations · Remote10h
- Head of Strategic Partnerships, TavilyUnited States · Director · RemoteSales · Director · RemoteUnited States15h
- General Manager, Data Center (New Build)Independence · Director · On-siteInformation Technology · Director · On-siteIndependence, United States1d
Let the right jobs find you
In your inbox every Wednesday and SaturdayPersonalised suggestions from verified career pages, matched to your role, location, level and skills.