Data Engineer II
New York, United States · Hybrid · Full-time
- Posted 3w ago
- From H1’s careers page
- Location
- New York, United States
- Work mode
- Hybrid
- Type
- Full-time
- Level
- Senior
- Experience
- 3+ years
- Department
- Data and Analytics
Opens the listing on jobs.lever.co
Let the right jobs find you
In your inbox every Wednesday and SaturdayPersonalised suggestions from verified career pages, matched to your role, location, level and skills.
About the role
WHAT YOU'LL DO AT H1
As a Data Engineer II on the H1DN team, you will build and operate the pipelines behind H1's clinical trials data. You'll work primarily in Python, PySpark, and SQL, and you'll work directly with clinical subject matter experts and Customer Success Managers to turn their domain knowledge into pipeline logic that holds up in production.
- Build and maintain the Python and PySpark pipelines behind the CTMS trial data pipeline intake workflows, including scoring and status logic.
- Develop the transformation logic that maps raw trial and customer data to H1's internal data models, handling diverse source formats including CSV, JSON, Parquet, and APIs.
- Write and tune SQL against large datasets to investigate data questions, validate pipeline output, and support analysis that clinical SMEs and customer-facing teams depend on.
- Turn around customer-driven changes quickly, scoping requests as they arrive, shipping changes that hold up under enterprise SLAs, and reworking logic as customer needs shift mid-flight.
- Partner with clinical SMEs to translate domain expertise into concrete data rules, then walk them through the results, explain what the pipeline did and why, and fold their feedback back into the logic.
- Build the data quality checks, validation logic, and reconciliation that let non-engineers trust pipeline output without reading the code.
- Participate in code reviews, maintaining a high bar for quality and adherence to engineering standards.
- Monitor and improve pipeline observability, contributing to alerting and dashboards that surface job health and data anomalies for both the team and internal users.
ABOUT YOU
You are a data engineer with a strong Python foundation and real distributed-processing experience. You're drawn to high-impact teams where the work is tangible: pipelines running, enterprise customers getting their data on time, clinical data that people make real decisions from. You're comfortable in an environment where recurring production runs and customer SLAs shape day-to-day priorities, and where a customer request can reorder your week. You'd rather sit down with a domain expert and understand why the data looks the way it does than build to a spec handed to you secondhand.
You bring experience:
- Building and shipping production data pipelines in Python, with an understanding of what makes them reliable and maintainable under real load
- Working with PySpark or a comparable distributed processing framework on datasets too large for a single machine
- Writing SQL well enough to answer hard questions about data, not just retrieve it
- Working in an operationally-driven environment where reliability and on-time delivery matter as much as new feature work
- Working directly with non-engineering partners, subject matter experts, analysts, or customer-facing teams, and communicating clearly about data with people who don't read code
- Holding a high bar in code review and expecting the same from those who review your work
- Identifying data problems early and seeing work through to resolution rather than handing it off
Skills they ask for
Pick one to see other roles that ask for it.
About H1
Provider and clinical data for better careH1 provides data and software that help life sciences, health plans, and digital health organizations identify medical providers, plan clinical trials, and support patient care.
See all 11 roles at H1More roles at H1
See all 11- Revenue Operations AnalystIndia · RemoteBusiness Operations · RemoteIndia1d
- Senior Forward Deployed Software EngineerNew York · Senior · HybridEngineering · Senior · HybridNew York, United States1w
- Data AnalystNew York · Senior · HybridData and Analytics · Senior · HybridNew York, United States3w
- Sr. Revenue Operations AnalystSenior · RemoteSales · Senior · Remote3w
Let the right jobs find you
In your inbox every Wednesday and SaturdayPersonalised suggestions from verified career pages, matched to your role, location, level and skills.