Machine Learning Engineer (Real-Time Speech Translation)
Washington D.C., United States · Hybrid · Full-time
- Posted 1w ago
- From LILT’s careers page
- Location
- Washington D.C., United States
- Work mode
- Hybrid
- Type
- Full-time
- Level
- Senior
- Experience
- 3+ years
- Department
- Engineering
Opens the listing on jobs.ashbyhq.com
Let the right jobs find you
In your inbox every Wednesday and SaturdayPersonalised suggestions from verified career pages, matched to your role, location, level and skills.
About the role
About LILT
AI is changing how the world communicates — and LILT is leading that transformation.
We're on a mission to make the world's information accessible to everyone, regardless of the language they speak. We use cutting-edge AI, machine translation, and human-in-the-loop expertise to translate content faster, more accurately, and more cost-effectively without compromising on brand, voice, or quality.
Role Summary
We are building a new live translation product. We are looking for an ML Engineer to build the real-time speech translation backend that powers it.
You will own the real-time speech translation backend end-to-end, from live audio input to translated output. You will build on LILT's production model serving platform (Ray Serve on GPU Kubernetes clusters) and our in-house adaptive machine translation models, working closely with the senior architects of that platform and with our language processing researchers. The ASR and MT models exist. Your job is to make them work together as a low-latency streaming system that holds up in production.
Key Responsibilities
-
Real-time pipeline architecture: Build and manage services for high-throughput, real-time audio and text streaming. Handle signal processing, session lifecycles, and concurrency management to ensure robust operation under load.
-
ML model integration: Integrate and serve streaming speech recognition and machine translation models, collaborating with research teams to ensure models operate within required latency budgets.
-
Quality and confidence workflows: Develop logic for model-based confidence scoring, routing segments for human intervention as needed, and broadcasting real-time updates and corrections to end-users.
-
Infrastructure and scale: Architect and scale production ML infrastructure on GPU-accelerated Kubernetes clusters. Implement batching, load balancing, and autoscaling strategies to maintain performance and cost-efficiency.
-
Latency engineering: Establish comprehensive instrumentation for real-time performance. Identify bottlenecks, optimize system throughput, and drive down end-to-end latency metrics to meet production standards.
-
Interface and API definition: Define technical contracts and interfaces for audio ingestion and downstream service integrations. Partner with frontend and platform engineering teams to maintain clean, robust integration points.
-
Collaboration and technical leadership: Drive cross-team alignment by defining clear API interfaces and technical contracts, facilitating effective communication between engineering and product teams to ensure seamless system integration.
Required Qualifications
-
BS or MS in Computer Science or a related field, or equivalent practical experience.
-
3+ years building production backend or ML serving systems in Python, including strong async programming (asyncio) skills.
-
Hands-on experience with real-time streaming transport: WebSocket or gRPC bidirectional streaming, session state, backpressure, and connection lifecycle handling.
-
Experience serving ML models in production on GPUs (Ray Serve, Triton, vLLM, or similar), with Docker and Kubernetes.
-
Experience integrating speech or NLP models into production systems, ideally streaming ASR (partial hypotheses, endpointing, VAD).
-
A latency-engineering mindset: you have profiled, instrumented, and optimized a real-time or low-latency system and can reason in per-stage budgets.
-
Effective use of AI coding agents (Claude Code, Codex, or similar) on top of fundamentals learned the hard way: you let agents do the typing, but you can debug, review, and reason about every line without them, and you know when not to trust them.
-
US citizenship and residence in the United States (contract requirement).
Preferred Qualifications
-
Ray Serve specifically, including streaming responses and model multiplexing.
-
Familiarity with simultaneous or incremental MT concepts (retranslation, prefix stability, wait-k policies).
-
Machine translation quality estimation (COMET/CometKiwi class models) or other confidence estimation in production.
-
Message brokers for real-time fan-out and state distribution (RabbitMQ or similar).
-
Streaming text-to-speech integration and time-to-first-audio optimization.
-
WebRTC and SFU concepts, or voice pipeline frameworks (LiveKit Agents, Pipecat).
-
Handling of CJK and other non-Latin text in NLP pipelines (our first languages are Japanese, Korean, and English).
-
Observability tooling (Datadog, Prometheus) for production ML systems.
Skills they ask for
Pick one to see other roles that ask for it.
About LILT
Enterprise AI translation platformLILT provides an enterprise AI translation platform for managing multilingual content and language workflows.
See all 32 roles at LILTMore roles at LILT
See all 32- Voice Talent Required - Hindi (female) - RemoteDelhi · RemoteCreative and Art Services · RemoteDelhi, India10h
- Linguist - Spanish (Puerto Rico) - RemoteSan Francisco · RemoteRemoteSan Francisco, United States15h
- Medical Translators - Hebrew - US-basedSan Francisco · Senior · RemoteSenior · RemoteSan Francisco, United States1d
- Medical Translators - Spanish (United States) - US-basedSan Francisco · Senior · RemoteSenior · RemoteSan Francisco, United States2d
Let the right jobs find you
In your inbox every Wednesday and SaturdayPersonalised suggestions from verified career pages, matched to your role, location, level and skills.