Staff MLOps Engineer
Palo Alto, United States · Hybrid · Full-time
- Posted 3h ago
- From AiDASH’s careers page
- Location
- Palo Alto, United States
- Work mode
- Hybrid
- Type
- Full-time
- Level
- Staff
- Experience
- 7+ years
- Department
- Engineering
Opens the listing on job-boards.greenhouse.io
Let the right jobs find you
In your inbox every Wednesday and SaturdayPersonalised suggestions from verified career pages, matched to your role, location, level and skills.
About the role
About AiDASH:
AiDASH is leading the PreventionFirst movement for electric utilities and transforming grid resilience through its platform that unifies vegetation, asset, storm, and wildfire intelligence. Powered by SatelliteFirst Inspection & Monitoring, AiDASH delivers comprehensive visibility across the entire grid. More than 200 customers trust AiDASH to keep the lights on, spend where it counts, and defend every decision.
The Role:
Our satellite and AI-powered products are changing how utilities see and protect the grid, and we're gearing up for aggressive customer growth over the next two years. We're building a brand-new Data Inference AI Pipeline team, and we're looking for a Staff MLOps Engineer to architect its operational backbone from the very first line of infrastructure code.
As the very first engineer on this team, you'll start with a blank canvas and a rare chance to build something that lasts. You'll shape how our models are deployed, how the pipeline scales and recovers under load, and what it costs to run. This is a hands-on, high-leverage individual contributor role, and you'll be the go-to technical authority on MLOps as the team grows. You'll report to the Director of Engineering for the Data Inference AI Pipeline team.
Our pending acquisition by Schneider Electric will accelerate our mission as we bring our combined offering to our joint customer base.
Location:
This is a hybrid role based in Palo Alto, CA, requiring 2 days per week in the office.
How you'll make an impact:
Deployment
- Own the model deployment pipeline end-to-end: packaging, versioning, rollout, and rollback for ML models moving from training to production
- Build and maintain deployment infrastructure on SageMaker (endpoints, batch transform, multi-model/multi-container hosting), and evaluate where alternatives such as self-hosted serving or other managed options are a better fit
- Set the CI/CD pattern for model releases, including automated testing, staged rollout, and canary/shadow deployments
Scalability & elasticity
- Design inference infrastructure that scales seamlessly with real-time and batch demand
- Right-size compute (CPU/GPU/Inferentia or equivalent) for each model and workload to balance latency and cost
- Load-test the pipeline and build capacity plans that stay ahead of customer growth milestones
Cost
- Own the cost-per-inference metric: instrument it, report on it, and drive it down as volume scales
- Build cost-awareness into architecture decisions, including autoscaling policies, spot/on-demand mix, and model size/quantization trade-offs
Failover & resilience
- Design for graceful degradation and failover across regions and availability zones for inference serving
- Define and test disaster-recovery procedures, and run regular failure-injection and chaos exercises
- Build monitoring, alerting, and on-call runbooks for model-serving infrastructure, covering drift, latency SLOs, error rates, and infra health
Cross-team leadership
- Partner with Data Science to bring new models to production on a platform built to support them
- Partner with the Web Application team where inference results power customer-facing features, keeping latency and reliability expectations aligned
- Serve as the technical authority on MLOps practices as the Data Inference AI Pipeline team grows
What success looks like in your first 6 months:
- Model deployment is a repeatable, automated process across models
- Inference infrastructure has proven it scales up and down smoothly under real load
- Cost-per-inference is measured, visible, and trending down as volume grows
- A tested failover and disaster-recovery plan is in place for the inference pipeline, with at least one failure-injection exercise completed
- Data Science can ship new models into production quickly and independently
Minimum Qualifications:
- 7-10+ years in software/ML engineering, including 3+ years focused specifically on MLOps or ML infrastructure
- Hands-on production experience with AWS SageMaker (or an equivalent managed ML platform) for model deployment and serving
- A proven track record designing systems for elasticity/autoscaling, cost optimization, and failover/resilience at production scale
- Strong expertise in containerization and orchestration (Docker, Kubernetes) and in CI/CD built for ML workloads, including model versioning and automated testing and rollout
- Experience instrumenting and owning cost and performance metrics for the infrastructure you're responsible for
- The ability to set technical direction as a Staff-level IC, building buy-in across Data Science and other engineering teams
- High ownership mindset, low ego, and a collaborative spirit, with a passion for building platforms that help others ship faster
Preferred Qualifications:
- Experience with large imagery, video, or 3D sensor and point-cloud data pipelines or other high-volume unstructured data at scale
- Experience with multimodal data (imagery, sensor/time-series, geospatial) beyond purely textual or tabular data
- Experience building MLOps practices from scratch on a new team
- Familiarity with model monitoring and drift-detection tooling, and GPU cost-optimization techniques (quantization, batching, Inferentia/Trainium or equivalent)
- A background in a regulated or enterprise B2B domain such as utilities, energy, infrastructure, healthcare, or finance
- Relevant AWS/Kubernetes certifications or open-source MLOps tooling contributions
What you'll love:
- Comprehensive Medical, Dental, and Vision Coverage: 100% coverage for employees and 80% for their spouses and children
- Health Reimbursement Account (HRA): 100% funded by AiDASH to cover medical deductibles
- 401(k) Plan: Begin contributing after two months of employment. Currently, no company match is offered
- Parental Leave: 16 weeks for primary caregivers and 4 weeks for secondary caregivers
- Generous Vacation Policy: Accrue 20 vacation days per year, plus an additional flex holiday
- Winter Break: From December 25th through January 1st
Compensation:
We offer a competitive annual pay range of $230,000 to $270,000 for this full-time position, which includes a base salary and bonus based on performance. This range reflects the anticipated annual pay range (base + bonus) for new hires.
We are proud to be an equal-opportunity employer and committed to embracing diversity and inclusion in our hiring practices.
Skills they ask for
Pick one to see other roles that ask for it.
About AiDASH
Satellite intelligence for infrastructure resilienceAiDASH provides satellite and AI-powered software and services for utilities and infrastructure operators to manage vegetation, wildfire risk, climate exposure, and asset inspections.
See all 11 roles at AiDASHMore roles at AiDASH
See all 11- SDET 2 - PlatformGurugramQuality AssuranceGurugram, India1d
- Manager - People & CultureBengaluru · Senior · On-siteHuman Resources · Senior · On-siteBengaluru, India1d
- Technical Support Engineer – Night Shift (Rotational)Gurugram · Entry Level · HybridCustomer Service · Entry Level · HybridGurugram, India3d
- Senior Technical Project ManagerPalo Alto · Senior · HybridProject and Program Management · Senior · HybridPalo Alto, United States4d
Let the right jobs find you
In your inbox every Wednesday and SaturdayPersonalised suggestions from verified career pages, matched to your role, location, level and skills.