Lead Support Engineer
Los Angeles, United States · Full-time
- Posted 4d ago
- From WPP Media’s careers page
- Location
- Los Angeles, United States
- Type
- Full-time
- Level
- Lead
- Department
- Information Technology
This role is no longer on WPP Media’s careers page. See 76 open roles
Let the right jobs find you
In your inbox every Wednesday and SaturdayPersonalised suggestions from verified career pages, matched to your role, location, level and skills.
About the role
Role Summary and Impact
This role is within WPP Media, where you will be instrumental in owning production stability, observability, and system health for a mission-critical global campaign governance and compliance platform. As Lead Support Engineer, you will provide advanced L2+ and L3-oriented support for a system running natively on Google Cloud Platform (GCP). You will investigate complex production issues, implement minor code-level fixes in Python, lead root cause analysis, and improve the reliability of distributed systems. Working closely with the EMEA-based Engineering Tech Lead, Product Owner, and Quality Assurance partners, you will connect production insights with technical roadmap priorities.
The opportunity combines hands-on Site Reliability Engineering (SRE) with technical squad leadership. A foundational monitoring and alerting setup is already in place, giving you the platform to evaluate the current approach, define a clear observability direction, and strengthen system health monitoring across the environment. You will automate runbooks, reduce manual operational effort, and help shape the future support squad for a dedicated product used across global advertising campaigns.
Key Responsibilities
- Own production stability, system health, and observability for a mission-critical campaign governance and compliance platform running on GCP.
- Diagnose and resolve complex, intermittent, and high-priority incidents across application, database, infrastructure, and networking layers.
- Read, debug, and implement minor fixes and patches directly within the existing Python production codebase.
- Define and advance the observability strategy using GCP Cloud Logging, Cloud Monitoring, Error Reporting, Prometheus, PromQL, and suitable service level indicators and objectives.
- Lead incident response, post-incident reviews, and end-to-end root cause analysis, partnering with the EMEA Engineering team on permanent remediation.
- Monitor execution flows, performance, data refreshes, automated checks, and recurring compliance reporting to ensure reliable system operation.
- Identify opportunities to automate manual support activity, develop and maintain runbooks, and reduce operational toil.
- Provide technical leadership to the L2+ support squad through coaching, knowledge sharing, prioritization, and effective handovers.
- Partner with Product, Engineering, and Quality Assurance teams to assess operational risk, improve release and change management, and influence technical roadmap decisions.
Skills and Experience
- Advanced education in computer science, software engineering, information technology, or a related technical discipline, or equivalent practical experience.
- Strong, current Python proficiency, including the ability to read, debug, and implement fixes directly within a shared production codebase.
- Deep, hands-on experience supporting enterprise systems on Google Cloud Platform, including Compute Engine, Google Kubernetes Engine, Cloud SQL, BigQuery, Pub/Sub, Firestore, Cloud Functions, Identity and Access Management, and virtual private cloud networking.
- Complex production troubleshooting and root cause analysis experience across distributed systems, including leadership of post-incident reviews.
- Strong experience with Docker and Kubernetes, particularly deploying, managing, and troubleshooting applications running in Google Kubernetes Engine.
- High proficiency with GCP Cloud Logging, Cloud Monitoring, Error Reporting, Prometheus, and PromQL, including establishing service level indicators and service level objectives.
- Experience using Terraform for infrastructure as code, tracing infrastructure-level issues, and identifying improvements to infrastructure management.
- Strong Python and Bash scripting skills for automating support tasks, parsing logs, developing operational tools, and reducing manual effort.
- Solid querying and troubleshooting experience across Firestore, BigQuery, PostgreSQL, and MySQL.
- Experience leading, mentoring, or serving as a technical lead for a squad or technical team, along with strong knowledge of Incident, Problem, and Change Management practices and the software development lifecycle.
Skills they ask for
Pick one to see other roles that ask for it.
About WPP Media
Intelligent growth for the AI eraWPP Media provides media and marketing services and technology to help clients drive growth.
See all 77 roles at WPP MediaMore roles at WPP Media
See all 77- Lead Support EngineerLos Angeles · Lead · HybridEngineering · Lead · HybridLos Angeles, United States6h
- Manager, PlanningNew York · HybridMarketing · HybridNew York, United States6h
- Manager, Commerce, IndiaGurugramMarketingGurugram, India12h
- Vice President, Strategy, IndiaGurugram · Executive · HybridMarketing · Executive · HybridGurugram, India13h
Let the right jobs find you
In your inbox every Wednesday and SaturdayPersonalised suggestions from verified career pages, matched to your role, location, level and skills.