Data Engineer II
Fam (FamApp)
Full Time3+ yearsPosted 22 days ago
Let the right jobs find you
Get personalised suggestions from verified company career pages, matched to your role, location, level, and skills.
Overview
Position Type
Full Time
Experience
3+ years
Job Description
About the role:
We are looking for a Data Engineer II (SDE-2) to join our data team. The ideal candidate will be a play a key role to develop of high performant and scalable Data Lake-house, moving us toward a world of sub-minute data latency and unified batch/streaming compute. This is an engineering-heavy role where you will manage complex CDC flows, optimize distributed query engines and leverage AI to accelerate our development lifecycle.
Must Have Requirements-
- Experience: 3–5 years in Data Engineering, specifically with distributed systems and cloud-native architectures.
- Coding: Expert-level Python/PySpark and SQL. Familiarity with Go/Java/Scala is a plus
- Infrastructure: Hands-on experience with AWS (S3, EKS, MSK) and Infrastructure-as-Code.
- Orchestration: Experience with Airflow or Temporal for complex workflow management.
- AI-Native: Proficiency in using AI tools (Claude, Codex, Copilot) to write, test, and document code efficiently.
- Systems Thinking: Ability to explain the trade-offs between different storage formats and processing frameworks.
- Tech Execution : Drive key tech initiatives by preparing TRD and actively involve in design reviews.
- Domain Modelling - Should be hands on in designing Domain models for OLAP like Fact, Dimension, Cumulative, types of SCD’s and OBT pattern tables.
- Self Starter - Lead the team technically and bring in new ideas to contribute to the growth of the charter.
- Stakeholder Interaction - Interact with the Product & Key Stakeholders & help them by adding value to the business workflow with data & analytics.
Good to have -
- Real-time CDC: Ownership of high-throughput ingestion from RDBMS to Lakehouse using Debezium, PeerDB.
- Lakehouse Architecture: Designing and optimizing table formats (Iceberg, Delta, Hudi) for both performance and storage efficiency.
- Unified Compute: Developing robust ETL/ELT frameworks in PySpark and Flink (handling both batch and streaming workloads).
- Infrastructure & Ops: Managing data workloads on AWS (EMR, EKS, MSK, S3) and automating everything via Gitlab/Github Actions.
- Query & BI: Tuning Trino or Clickhouse to power real-time dashboards in Metabase, Superset, and PowerBI.
Our Tech Stack-
- Ingestion & CDC: OLake and PeerDB for near real-time sync from production systems into Alchemy; Debezium/Kafka for CDC-heavy use cases, with support for Kafka and S3-based sources.
- Lakehouse / Storage: Alchemy on Apache Iceberg with S3 as the data lake storage layer and AWS Glue Catalog for metadata; exposure to Delta/Hudi is a plus.
- Processing & Compute: PySpark on EMR/EKS for batch and streaming workloads; Flink and Spark Structured Streaming fundamentals for low-latency pipelines.
- Streaming Platform: MSK / Kafka for event-driven ingestion, CDC propagation, replay, backfills, and operational monitoring through Kafka UI.
- Query & Serving Layer: Trino over Alchemy/Iceberg for lakehouse analytics, ClickHouse for high-throughput operational and real-time dashboards, and BigQuery exposure where applicable.
- Workflow Orchestration: Airflow for scheduled data pipelines, DQ/reconciliation DAGs, backfills, and SLA-driven jobs; Temporal for durable workflow execution in ingestion services.
- Data Quality & Governance: DQ checks, freshness/SLA monitoring, source-to-lake reconciliation, deduplication, schema evolution handling, and cataloging/lineage through OpenMetadata.
- Infrastructure & DevOps: AWS (S3, EKS, EMR, MSK), Kubernetes, Terraform/IaC, GitLab/GitHub Actions, observability via Grafana/CloudWatch, and production runbook discipline.
- BI & Analytics: Metabase, Superset, Tableau, and PowerBI for business dashboards; strong ability to model curated marts, fact/dimension tables, SCDs, and OBT patterns for stakeholder-facing analytics.