Research Engineer, Data Infrastructure (Language Modeling)
San Francisco, United States · On-site · Full-time
- Posted 2w ago
- From Cartesia’s careers page
- Location
- San Francisco, United States
- Work mode
- On-site
- Type
- Full-time
- Department
- Research and Development (R&D)
Opens the listing on jobs.ashbyhq.com
Let the right jobs find you
In your inbox every Wednesday and SaturdayPersonalised suggestions from verified career pages, matched to your role, location, level and skills.
About the role
About the Role
Data is the lifeblood of our models, and we are looking for a Research Engineer, Data Infrastructure to build the datasets and systems that power pretraining at Cartesia. In this role, you will write performant, scalable infrastructure to acquire, process, and curate massive datasets, and partner closely with research to optimize the characteristics and composition of data mixtures. Your work will directly shape the capabilities and quality of our foundational models.
Your Impact
-
Build and operate performant, scalable data processing infrastructure for acquiring, ingesting, and combining massive text datasets.
-
Design and operate scalable, high-throughput, and reproducible data pipelines — covering ingestion, preprocessing, filtering, deduplication, and augmentation.
-
Design and run ablation experiments to understand how data sources, processing choices, and mixture weights affect model quality.
-
Partner closely with research and infrastructure teams to co-design data loading, versioning, and experimentation pipelines.
-
Establish and enforce rigorous standards for data quality, with a tight feedback loop between dataset characteristics and model behavior.
-
Identify and source novel datasets; manage relationships and budgets with external data vendors and partners.
What You Bring
-
Hands-on experience with ML data infrastructure: training data pipelines, dataset versioning, large-scale data loading, and the interplay between data systems and model training and inference.
-
Strong modern engineering execution: clean, well-tested code, fluency with current tools, and a willingness to pick the right tool for the problem rather than defaulting to familiar patterns.
-
Familiarity with building and evaluating datasets for generative models and reasonable working knowledge of how they're trained and inference.
Nice-To-Haves
-
Experience with large-scale data processing using parallel infrastructure such as Ray, Spark, or Kubernetes.
-
Experience with pretraining language models.
Skills they ask for
Pick one to see other roles that ask for it.
About Cartesia
Voice AI that speaks and listens naturallyCartesia develops AI models and products for speech generation and real-time voice interaction.
See all 11 roles at CartesiaMore roles at Cartesia
See all 11- Strategic Partnerships Manager, Systems IntegratorsSan Francisco · On-siteBusiness Development · On-siteSan Francisco, United States2d
- Product RoleBengaluruProduct ManagementBengaluru, India1w
- Commercial Account ExecutiveSan Francisco · Senior · On-siteSales · Senior · On-siteSan Francisco, United States2w
- GTM Associate, InboundSan Francisco · Entry Level · On-siteEntry Level · On-siteSan Francisco, United States2w
Let the right jobs find you
In your inbox every Wednesday and SaturdayPersonalised suggestions from verified career pages, matched to your role, location, level and skills.