Senior Staff Machine Learning Engineer
Bengaluru, India · Full-time
- Posted 1w ago
- From Netradyne’s careers page
- Location
- Bengaluru, India
- Type
- Full-time
- Level
- Senior
- Experience
- 8+ years
- Department
- Data and Analytics
Apply on Netradyne’s site
Opens the listing on netradyne.com
Let the right jobs find you
In your inbox every Wednesday and SaturdayPersonalised suggestions from verified career pages, matched to your role, location, level and skills.
About the role
Job Responsibilities:
- Owning the architecture of large-scale cloud ML systems end to end — data ingestion and feature pipelines, training infrastructure, model serving, monitoring and retraining.
- Design, develop and deploy production ready scalable cloud solutions that utilize Gen-AI, agentic AI, DNN, Traditional ML models, data-driven rules and ETL pipelines.
- Architecting distributed, fault-tolerant services and data platforms that operate reliably at high throughput, with clear SLAs, observability and cost controls.
- Applying advanced statistical methods, machine learning and deep learning techniques to uncover trends, patterns, and anomalies in large-scale datasets.
- Creating robust frameworks and tools to automate and enhance data mining, labeling, model training, and validation processes for internal ML/DL initiatives.
- Setting engineering standards across teams — design review, testing strategy, CI/CD and release practice — and mentoring Staff and Senior engineers.
- Collaborating closely with cross-functional teams to identify and implement data-driven solutions addressing key business challenges.
- Conducting studies, setting up automation tools and frameworks, and regularly publishing internal and external KPI audits.
- Develop and maintain ROI models and frameworks to quantify the business impact of data science initiatives.
Requirements:
- B. Tech, M. Tech or PhD in Data Science, Computer Science, Electrical Engineering, Operations Research, Statistics, Mathematics or a related area.
- At least 8 years of working experience in machine learning, data science or a related domain, including 5+ years building and shipping production ML systems at scale.
- Demonstrated depth on both sides of the role: building distributed data and ETL pipelines, and training, tuning and deploying models in production.
- Strong large-scale software engineering fundamentals: distributed systems, concurrency, microservice and API design, caching, queueing, idempotency and failure handling.
- Proven experience designing and operating systems on public cloud at scale — AWS preferred (Kinesis, SQS, EKS, Lambda, Auto Scaling Groups, S3), including cost, capacity and reliability trade-offs.
- Strong foundational knowledge in Statistics, Probability Theory, Machine Learning and Gen-AI.
- Excellent programming skills – Python (required) and Java/Rust/C++ (desired), with strong fundamentals in object-oriented programming, algorithms, and data structures.
- Good understanding of database internals and schema design for relational (RDBMS) and non-relational (NoSQL) data stores, including the ability to write and reason about complex SQL.
- Experience with transformer architectures and large language models (LLMs), and with Gen-AI tools and workflows.
- Knowledge of best practices in software development, including version control, code review, automated testing, continuous integration and continuous delivery.
- Experience with observability and production operations — metrics, tracing, logging, alerting and incident response for services and ML pipelines.
- Proven ability to influence technical decisions beyond one's own team.
Desired skills:
- Agentic AI systems — tool use, planning, multi-agent orchestration, memory, guardrails and agent evaluation.
- Hands-on experience with the Claude Agent SDK, OpenAI Agents SDK and Model Context Protocol (MCP) servers and connectors.
- AI-native development practice — working effectively with coding agents, GitHub Copilot, Claude Code or similar, and setting team conventions for their use.
- Test-Driven Development (TDD) and Spec-Driven Development (SDD); designing specs and evals that agents and humans can both work against.
- LLMOps: prompt and context management, retrieval-augmented generation, model routing, caching, token cost optimisation and offline/online eval harnesses.
- Infrastructure as code and container orchestration — Terraform, Kubernetes, Helm; multi-region and blue-green or canary deployment patterns.
- Streaming and service technologies such as Kafka streams, Queues, Rest API and gRPC systems.
- Tools: FastAPI, MLFlow, Huggingface pipelines, LangGraph, OpenAI, Anthropic API.
- Experience with MLOps tools and practices for continuous deployment and monitoring of AI models.
- Experience with data visualization tools like Tableau, Grafana, Plotly-Dash.
Skills they ask for
Pick one to see other roles that ask for it.
About Netradyne
More roles at Netradyne
See all 14- Manager Revenue OperationsBengaluruBusiness OperationsBengaluru, India12h
- Associate Sales Enablement SpecialistBengaluruSalesBengaluru, India1d
- Senior QA Engineer - Vehicle Networks and TelematicsBengaluru · SeniorQuality Assurance · SeniorBengaluru, India3d
- Associate Technical Program ManagerBengaluru · Entry LevelProduct Management · Entry LevelBengaluru, India1w
Share this role
Let the right jobs find you
In your inbox every Wednesday and SaturdayPersonalised suggestions from verified career pages, matched to your role, location, level and skills.