LLMOps Engineer

Gnani.ai

Full Time3+ yearsPosted 4 days ago

Let the right jobs find you

Get personalised suggestions from verified company career pages, matched to your role, location, level, and skills.

Overview

Position Type

Full Time

Experience

3+ years

Job Description

Job Description

Responsibilities

  • Own the production LLM serving stack across vLLM, SGLang, and NVIDIA Dynamo
  • Tune the serving path end to end
  • Design and operate disaggregated prefill/decode and multi-node deployments
  • Run inference on Kubernetes with autoscaling, rolling model updates, canary releases, and tested rollback paths
  • Instrument everything that matters
  • Build and maintain internal serving APIs, model registries, and deployment tooling

Model Optimization

  • Quantization: FP8, INT8, and INT4 weight and activation quantization plus KV cache quantization
  • Speculative decoding: draft-model, EAGLE/Medusa-style, and n-gram approaches
  • Distillation: build smaller task-specific students from larger teachers for latency-critical paths
  • Pruning and adapters: structured pruning, LoRA/adapter serving, and multi-adapter batching for per-deployment specialization
  • Compilation: torch.compile, CUDA graphs, and TensorRT-LLM engine builds

Kernel and Low-Level Performance

  • Profile with Nsight Systems/Compute and the PyTorch profiler
  • Write and tune custom kernels in CUDA and Triton
  • Eliminate host-device synchronization stalls, dynamic-shape recompilation, and unoverlapped collectives in distributed serving
  • Work fluently with attention and MoE kernel libraries (FlashAttention, FlashInfer, CUTLASS/cuBLAS, expert-parallel dispatch libraries) and know when to use them versus write your own
  • Reason from first principles about arithmetic intensity, memory bandwidth, and roofline limits before reaching for a tool

Evaluation, Reliability, and Cost

  • Own the optimization regression gate
  • Build load-testing harnesses that replay realistic traffic
  • Run capacity planning and cost modeling
  • Carry on-call for inference services, write runbooks, and lead blameless postmortems on latency and availability incidents
  • Work closely with the training and post-training teams so that serving constraints inform model architecture decisions early, not after the checkpoint lands

Must Have

  • 3 to 6 years in ML infrastructure, model serving, or performance engineering, with at least 2 years specifically on LLM inference in production
  • Hands-on production experience with at least two of vLLM, SGLang, TensorRT-LLM, or NVIDIA Dynamo, and the ability to explain how their schedulers and KV cache designs differ
  • Demonstrated quantization work on real models (FP8/INT8/INT4) with measured accuracy impact and a defensible calibration methodology
  • End-to-end experience with at least one of speculative decoding, distillation, or structured pruning
  • Strong Python; comfortable reading and patching PyTorch and serving-engine source code
  • CUDA fundamentals: memory hierarchy, occupancy, kernel launch overhead, roofline reasoning; able to read and interpret a profiler trace
  • Multi-GPU serving: tensor and pipeline parallelism, NCCL basics, and debugging distributed hangs and stragglers
  • Kubernetes, Docker, and GPU scheduling in production; observability with Prometheus/Grafana or equivalent
  • Ability to quantify your own work in latency percentiles, throughput per GPU, and cost per million tokens

Good to Have

  • Real-time speech inference: streaming ASR/TTS, sub-second time-to-first-audio budgets, barge-in and turn-taking constraints
  • Serving hybrid-architecture models (state-space/Mamba blocks combined with attention) and understanding their distinct state and cache management
  • TensorRT-LLM engine building, NVIDIA NIM, or Triton Inference Server in production
  • Experience with Indic or other multilingual and code-mixed models, and the evaluation discipline that comes with them
  • Ray Serve, KServe, LLM gateways, or semantic and prefix cache layers
  • Enough training-side exposure (Megatron, NeMo, DeepSpeed, FSDP) to work fluently with the pretraining team
  • Public benchmarks, blog posts, or talks on inference optimization

What You Will Work On

You will work on in-house foundational models — not third-party API endpoints — running on large NVIDIA GPU clusters and serving live enterprise and government-scale voice traffic. The models are ours, so the whole stack is open to you: you can change the serving engine, the quantization recipe, the kernel, or the model itself, and you will often need to change more than one to hit a target.

The constraints are real and the feedback loop is fast. A latency regression is heard by callers within minutes. A successful optimization shows up directly in infrastructure spend. You will have access to proprietary speech and language data, in-house training and evaluation infrastructure, a close engineering partnership with NVIDIA, and a team that has been building Indian-language AI since 2016.

This role has a clear path to owning the inference platform and its technical direction, or to a deeper specialization in performance and kernel engineering, depending on where your strength lies.

Required Skills

LlmAgentic LlmTensort LlmRag Agentic Systems V Llm Sg Lang Vector Databases

About the Company

Gnani.ai

Bengaluru, India

Share This Job