Principal Data Engineer
Pune, India
- Posted 1mo ago
- From HG Insights’s careers page
- Location
- Pune, India
- Level
- Principal
- Experience
- 15+ years
- Department
- Engineering
Opens the listing on hginsights.com
Let the right jobs find you
In your inbox every Wednesday and SaturdayPersonalised suggestions from verified career pages, matched to your role, location, level and skills.
About the role
About the Role
This is a senior individual contributor role on the team that builds and runs HG's data platform: everything between a raw vendor file landing in our lake and a finished dataset arriving in a customer's hands. It includes the transformation stack that turns messy external signals into product-grade company attributes, the matching systems that decide which company a piece of evidence belongs to and the internal applications our analysts use to curate and correct the result.
What you’ll do
- Own the architecture of the data platform end to end, and make the calls on build vs. buy, batch vs. incremental, and where each workload belongs.
- Build the quality layer that catches silent failure: contracts at every handoff, freshness and completeness monitoring, and pipelines that quarantine bad data rather than publish it.
- Improve entity matching and attribution — and make match decisions explainable to the customers who depend on them.
- Make our release cycle boring. Recurring deliveries should not depend on people watching them.
- Treat compute cost as a first-class engineering metric across Spark, orchestration, and the serving tier.
What we are looking for
- 15+ years building production grade data engineering systems, with a minimum of 5 years in a staff or principal scope.
- Deep Spark and Databricks expertise at multi-terabyte scale, including the operational side of lakehouse table formats — merges, compaction, small files, schema evolution.
- Strong SQL and dimensional modelling, and the judgment to know when to break the rules.
- Experience with performant and scalable OLTP setups (MySQL/Postgres).
- Airflow orchestration (DAGs, operators, sensors) and integration with Spark/Databricks.
- Proven experience in AWS ecosystems (EC2, S3, EMR).
- Hands-on entity resolution, fuzzy matching, or record linkage at scale. This is central to what we do.
- Python/Scala/Java, with real software discipline: testing, CI/CD, code review, infrastructure as code is a plus.
- Experience in Docker - Kubernetes, Terraform is a plus.
- Experience in integrating AI first implementations in traditional data engineering setups will greatly help in shaping our future designs.
- Experience with machine learning pipelines (Spark MLlib, Databricks ML) for predictive analytics.
- Knowledge of data governance frameworks and compliance standards (GDPR, CCPA).
Skills they ask for
Pick one to see other roles that ask for it.
About HG Insights
Revenue intelligence for B2B teamsHG Insights provides B2B market and technology intelligence to help organizations plan and execute go-to-market work.
See all 7 roles at HG InsightsMore roles at HG Insights
See all 7- Review Generation Team Lead & Community ManagerUnited States · Senior · RemoteMarketing · Senior · RemoteUnited States1w
- Marketing Operations ManagerPune · SeniorMarketing · SeniorPune, India3w
- Customer Success Manager IIPune · SeniorCustomer Service · SeniorPune, India3w
- Principal UX designerPune · PrincipalProduct Management · PrincipalPune, India1mo
Let the right jobs find you
In your inbox every Wednesday and SaturdayPersonalised suggestions from verified career pages, matched to your role, location, level and skills.