Data Infra Engineer
Waymo
Full Time4+ yearsPosted 10 days ago
Let the right jobs find you
Get personalised suggestions from verified company career pages, matched to your role, location, level, and skills.
Overview
Position Type
Full Time
Experience
4+ years
Job Description
Roles and Responsibilities:
- Design, develop, and maintain high-throughput distributed backend components and services (C++, Flume/Apache Spark/Flink/Beam) powering Waymo's Commercialization Data Lake.
- Own the scalable ingestion of data from raw system telemetry and RPC logs, processing petabytes of logs into reliable, daily canonical base tables.
- Build, harden, and democratize Data Pipeline Services, offering standard templates and self-service gateways for downstream consumption.
- Implement foundational compliance, privacy, and security infrastructure (e.g., GDPR wipeout, ACL management, BCID enforcement, and identity proxies).
- Establish rigorous operational standards, guarding pipelines with regression tests, staging environments, Canary/Progressive rollouts to defend high SLO.
- Collaborate with Data Engineers and Product Data Scientists to set global platform consistency, dependency tracing, and linear DAG architecture guidelines.
You have:
- BS/MS in Computer Science, Computer Engineering, or equivalent.
- 4+ years of professional backend software engineering experience, focused heavily on distributed systems, databases, and production monitoring.
- Excellent understanding of data structures, algorithms, and fundamental software design principles.
- Advanced proficiency in C++ and SQL applied to high-throughput batch and streaming processing pipelines.
- Prior experience running data pipelines through hardened release setups (Canarying, Staging environments, Automated Rollbacks).
We prefer:
- Extensive experience with distributed processing frameworks (e.g., Spark, Flink, Beam, Hadoop, Flume, or equivalent large-scale processing engines).
- Experience in building and maintaining Data Warehouse / Data Lake systems.
- Experience baking Data Governance, Privacy, or Compliance layers (e.g., encryption, GDPR, RBAC/ACLs) directly into core infrastructure.
- High proficiency with Google-internal pipeline/data stacks (Flume, SQLP, Rapid, CAS, F1).