Lead Data Engineer
India · Remote · Full-time
- Posted 2w ago
- From Algoworks’s careers page
- Location
- India
- Work mode
- Remote
- Type
- Full-time
- Level
- Senior
- Experience
- 10+ years
- Department
- Data and Analytics
Apply on Algoworks’s site
Opens the listing on algoworks.keka.com
Let the right jobs find you
In your inbox every Wednesday and SaturdayPersonalised suggestions from verified career pages, matched to your role, location, level and skills.
About the role
Role:
Lead Data Engineer
Location:
India, Remote
Experience:
10+ Years
Role overview
We are seeking a hands-on Senior Data Engineer with strong expertise in Azure Databricks and Azure Data Factory to build, optimize, and maintain scalable enterprise data pipelines.
Key responsibilities:
1. Pipeline Development
- Build and maintain scalable data pipelines using Azure Databricks and Azure Data Factory.
- Implement ingestion and transformation logic across Bronze and Silver data layers.
- Develop batch and incremental data-processing patterns.
- Design reliable and reusable pipeline components for enterprise workloads.
- Monitor and troubleshoot pipeline execution and data-processing issues.
2. Curated Layer & Delta Lake Development
- Implement hydration, merge, and upsert logic using Delta Lake.
- Build and maintain curated datasets aligned with data quality and business requirements.
- Handle late-arriving data and incremental updates.
- Implement reliable data transformation and reconciliation processes.
- Ensure curated datasets are optimized for downstream consumption.
3. Performance & Storage Optimization
- Optimize Delta Lake tables for performance and cost efficiency.
- Select and tune appropriate storage formats such as Parquet and Delta.
- Apply partitioning, compaction, and file-sizing strategies.
- Tune Spark jobs for large-scale distributed data processing.
- Identify and resolve performance bottlenecks across data pipelines and storage layers.
4. Downstream & DWH Collaboration
- Work closely with DWH and reporting teams to support downstream data consumption.
- Provide optimized datasets for reporting and analytical workloads.
- Support data validation and reconciliation with Gold-layer outputs.
- Collaborate with downstream teams to understand data requirements and optimize delivery.
- Ensure consistency and reliability of data consumed by reporting and analytics platforms.
5. Engineering Best Practices
- Implement basic CI/CD practices for data pipelines.
- Follow coding standards, documentation, and version-control practices.
- Maintain reusable, scalable, and maintainable pipeline code.
- Support production troubleshooting and performance tuning.
- Participate in Agile delivery processes and technical discussions.
6. Data Quality & Production Support
- Implement data validation and quality checks across ingestion and transformation processes.
- Investigate data discrepancies and pipeline failures.
- Perform root-cause analysis and implement corrective actions.
- Support production deployments and resolve data-processing issues.
- Maintain reliability and consistency across enterprise data pipelines.
Skills they ask for
Pick one to see other roles that ask for it.
About Algoworks
AI and digital engineering servicesAlgoworks provides AI, digital engineering, Salesforce, DevOps and mobile development services to businesses.
See all 28 roles at AlgoworksMore roles at Algoworks
See all 28- Software Engineer – Quality Assurance AutomationNoidaQuality AssuranceNoida, India23h
- Integration EngineerNoida · SeniorSoftware Development · SeniorNoida, India1d
- Backend Java DeveloperIndia · SeniorSoftware Development · SeniorIndia2d
- Salesforce Scrum MasterNoida · SeniorInformation Technology · SeniorNoida, India2d
Share this role
Let the right jobs find you
In your inbox every Wednesday and SaturdayPersonalised suggestions from verified career pages, matched to your role, location, level and skills.