LLMOps Engineer
Gnani.ai
Let the right jobs find you
Get personalised suggestions from verified company career pages, matched to your role, location, level, and skills.
Overview
Position Type
Full Time
Experience
6+ years
Job Description
Job Description
Responsibilities
- Own the production LLM serving stack across vLLM, SGLang, and NVIDIA Dynamo
- Tune the serving path end to end
- Design and operate disaggregated prefill/decode and multi-node deployments
- Run inference on Kubernetes with autoscaling, rolling model updates, canary releases, and tested rollback paths
- Instrument everything that matters
- Build and maintain internal serving APIs, model registries, and deployment tooling
Model Optimization
- Quantization: FP8, INT8, and INT4 weight and activation quantization plus KV cache quantization
- Speculative decoding: draft-model, EAGLE/Medusa-style, and n-gram approaches
- Distillation: build smaller task-specific students from larger teachers for latency-critical paths
- Pruning and adapters: structured pruning, LoRA/adapter serving, and multi-adapter batching for per-deployment specialization
- Compilation: torch.compile, CUDA graphs, and TensorRT-LLM engine builds
Kernel and Low-Level Performance
- Profile with Nsight Systems/Compute and the PyTorch profiler
- Write and tune custom kernels in CUDA and Triton
- Eliminate host-device synchronization stalls, dynamic-shape recompilation, and unoverlapped collectives in distributed serving
- Work fluently with attention and MoE kernel libraries (FlashAttention, FlashInfer, CUTLASS/cuBLAS, expert-parallel dispatch libraries) and know when to use them versus write your own
Evaluation, Reliability, and Cost
- Own the optimization regression gate
- Build load-testing harnesses that replay realistic traffic
- Run capacity planning and cost modeling
- Carry on-call for inference services, write runbooks, and lead blameless postmortems on latency and availability incidents
- Work closely with the training and post-training teams so that serving constraints inform model architecture decisions early, not after the checkpoint lands
Must Have
- 6 to 10 years total experience, with 3+ years owning large-scale LLM or speech inference in production
- Has owned an inference platform end to end: architecture, SLOs, capacity, cost, deployment safety, and on-call
- Has written or substantially tuned custom CUDA or Triton kernels that shipped to production
- MoE serving at scale: expert parallelism, routing load imbalance, dispatch kernels, and the failure modes specific to sparse models
- Deep profiling ability
- Track record running an accuracy-preserving optimization program (quantization plus speculative decoding plus distillation) behind real evaluation gates
- Sets technical direction, mentors engineers, and can make a defensible build-versus-adopt call on serving infrastructure
- Upstream contributions to open-source serving or kernel projects (vLLM, SGLang, FlashInfer, TensorRT-LLM, or similar) — a strong plus, and close to an expectation at this level
Good to Have
- Real-time speech inference: streaming ASR/TTS, sub-second time-to-first-audio budgets, barge-in and turn-taking constraints
- Serving hybrid-architecture models (state-space/Mamba blocks combined with attention) and understanding their distinct state and cache management
- TensorRT-LLM engine building, NVIDIA NIM, or Triton Inference Server in production
- Experience with Indic or other multilingual and code-mixed models, and the evaluation discipline that comes with them
- Ray Serve, KServe, LLM gateways, or semantic and prefix cache layers
- Enough training-side exposure (Megatron, NeMo, DeepSpeed, FSDP) to work fluently with the pretraining team
- Public benchmarks, blog posts, or talks on inference optimization
What You Will Work On
You will work on in-house foundational models — not third-party API endpoints — running on large NVIDIA GPU clusters and serving live enterprise and government-scale voice traffic. The models are ours, so the whole stack is open to you: you can change the serving engine, the quantization recipe, the kernel, or the model itself, and you will often need to change more than one to hit a target.
The constraints are real and the feedback loop is fast. A latency regression is heard by callers within minutes. A successful optimization shows up directly in infrastructure spend. You will have access to proprietary speech and language data, in-house training and evaluation infrastructure, a close engineering partnership with NVIDIA, and a team that has been building Indian-language AI since 2016.
This role has a clear path to owning the inference platform and its technical direction, or to a deeper specialization in performance and kernel engineering, depending on where your strength lies.