Machine Learning Research Scientist, Evaluations
San Francisco, United States · Full-time
- Posted 1mo ago
- From Scale AI’s careers page
- Location
- San Francisco, United States
- Type
- Full-time
- Department
- Research and Development (R&D)
Apply on Scale AI’s site
Opens the listing on job-boards.greenhouse.io
Let the right jobs find you
In your inbox every Wednesday and SaturdayPersonalised suggestions from verified career pages, matched to your role, location, level and skills.
About the role
Roles and Responsibilities:
- Analyze model behavior to identify, characterize, and diagnose failure modes in frontier LLMs and Agents. You’ll identify everything from capability gaps and reasoning errors to robustness and alignment issues, all focusing on RCA.
- Design and build benchmarks and evaluation methods that measure LLM capabilities in both text and multimodal modalities.
- Apply post-training expertise (SFT, RLHF, reward modeling) to connect observed failures to the data and training interventions that address them.
- Publish research findings in top-tier AI conferences.
Ideally you’d have:
- Ph.D. or Master's degree in Computer Science, Machine Learning, AI, or a related field.
- Deep understanding of deep learning, reinforcement learning, and large-scale model fine-tuning.
- Experience with post-training techniques such as RLHF, preference modeling, or instruction tuning, and with LLM evaluation or benchmark development.
- Excellent written and verbal communication skills.
- Published research in areas of machine learning at major conferences (NeurIPS, ICML, ICLR, ACL, EMNLP, CVPR, etc.) and/or journals.
- Previous experience in a customer facing role.
Skills they ask for
Pick one to see other roles that ask for it.
About Scale AI
Reliable AI systems for critical decisionsScale AI develops data and AI systems used to support reliable AI and critical decisions across the AI stack.
See all 35 roles at Scale AIMore roles at Scale AI
See all 35- Systems Engineering Integration & Test Lead, Public SectorWashington · LeadEngineering · LeadWashington, United States10h
- Enterprise AI Development StrategistNew York · Entry LevelSales · Entry LevelNew York, United States2d
- Product Operations Lead, Generative AISan Francisco · SeniorBusiness Operations · SeniorSan Francisco, United States1w
- Revenue Operations ManagerSan Francisco · SeniorBusiness Operations · SeniorSan Francisco, United States1w
Share this role
Let the right jobs find you
In your inbox every Wednesday and SaturdayPersonalised suggestions from verified career pages, matched to your role, location, level and skills.