Lead Data Engineer
Bengaluru, India · Hybrid · Full-time
- Posted 1w ago
- From Relanto’s careers page
- Location
- Bengaluru, India
- Work mode
- Hybrid
- Type
- Full-time
- Level
- Lead
- Experience
- 6+ years
- Department
- Data and Analytics
Apply on Relanto’s site
Opens the listing on relanto.keka.com
Let the right jobs find you
In your inbox every Wednesday and SaturdayPersonalised suggestions from verified career pages, matched to your role, location, level and skills.
About the role
Role Summary
We are looking for a Lead Data Engineer with strong hands-on experience in real-time data streaming, event processing, CDC, data integration, and modern data engineering.
Key Responsibilities
- Real-Time Data Engineering
- Design, develop, and maintain real-time data pipelines using Apache Flink and Apache Kafka.
- Develop production-grade streaming applications for high-volume and low-latency workloads.
- Implement data transformation, filtering, enrichment, aggregation, and event processing.
- Build reliable event-processing pipelines with appropriate error handling and recovery mechanisms.
- Consume and publish events across Kafka topics.
- Implement appropriate partitioning, consumer groups, offsets, and delivery mechanisms.
- Troubleshoot streaming pipeline failures and performance issues.
Apache Flink
- Develop and maintain Apache Flink jobs for real-time data processing.
- Implement:
- Stream transformations
- Filtering
- Mapping
- Aggregations
- Joins
- Windows
- Event-time processing
- Watermarks
- State management
- Implement Flink checkpointing and recovery mechanisms.
- Optimize Flink jobs for performance, scalability, and resource utilization.
- Monitor Flink jobs for latency, throughput, failures, and resource consumption.
- Troubleshoot state, checkpointing, backpressure, and processing issues.
Kafka
- Develop Kafka-based ingestion and streaming pipelines.
- Create and manage Kafka topics and event streams.
- Work with partitions, offsets, consumer groups, replication, and retention.
- Implement reliable producer and consumer applications.
- Handle message ordering, retries, duplicate events, and replay scenarios.
- Monitor Kafka performance and troubleshoot consumer lag and throughput issues.
- Work with Kafka schemas and serialization formats.
CDC & Debezium
- Build CDC-based ingestion pipelines using Debezium.
- Configure and maintain Debezium connectors.
- Capture source-system inserts, updates, and deletes.
- Publish CDC events into Kafka.
- Handle initial snapshots and incremental CDC processing.
- Manage schema evolution and changes in source systems.
- Implement data reconciliation and consistency checks.
- Troubleshoot CDC failures and source-to-target data issues.
Data Orchestration
- Develop and maintain data workflows using Apache Airflow or equivalent orchestration frameworks.
- Build reusable DAGs for:
- Data ingestion
- CDC workflows
- Data validation
- Flink job execution
- Data transformation
- ClickHouse loading
- Downstream integrations
- Implement workflow dependencies, scheduling, retries, backfills, SLAs, and alerting.
- Integrate Airflow workflows with Kafka, Flink, Debezium, ClickHouse, APIs, and cloud services.
- Monitor workflow execution and troubleshoot failures.
- Develop reusable operators, sensors, and workflow components where appropriate.
- Use event-driven triggers where real-time workflows require them.
ClickHouse & Analytical Data
- Integrate streaming data pipelines with ClickHouse.
- Design efficient analytical data models.
- Develop and optimize SQL queries.
- Implement appropriate partitioning, sorting, indexing, and retention strategies.
- Optimize data ingestion and query performance.
- Support downstream analytical use cases, dashboards, and reporting requirements.
Data Quality & Reliability
- Implement data validation and quality checks throughout the pipeline.
- Build reconciliation mechanisms between source and target systems.
- Monitor data freshness, completeness, accuracy, and consistency.
- Implement error handling, retry, replay, and recovery mechanisms.
- Establish logging and observability for critical pipelines.
- Support incident investigation and root-cause analysis.
Integration & APIs
- Integrate streaming and analytical data with APIs, endpoints, dashboards, and downstream applications.
- Develop data interfaces and integration components.
- Work with application teams to define data contracts and integration requirements.
- Support future integrations and additional data consumers.
Engineering Practices
- Follow modern software engineering practices around:
- Git
- Code reviews
- Unit testing
- Integration testing
- CI/CD
- Logging
- Monitoring
- Documentation
- Develop reusable and maintainable data engineering components.
- Participate in technical design discussions and architecture reviews.
- Mentor other Data Engineers and contribute to engineering standards.
Skills they ask for
Pick one to see other roles that ask for it.
About Relanto
AI-informed business consulting and transformationRelanto is a business consulting and advisory firm that applies AI, data, and technology to planning and transformation across industries.
See all 50 roles at RelantoMore roles at Relanto
See all 50- Lead Snowflake DBT DeveloperBengaluru · LeadData and Analytics · LeadBengaluru, India2d
- Inside Sales ExecutiveBengaluru · SeniorSales · SeniorBengaluru, India2d
- Technical Program ManagerBengaluru · SeniorProject and Program Management · SeniorBengaluru, India2d
- Lead Backend EngineerBengaluru · LeadSoftware Development · LeadBengaluru, India2d
Share this role
Let the right jobs find you
In your inbox every Wednesday and SaturdayPersonalised suggestions from verified career pages, matched to your role, location, level and skills.