Job Title: L2 Data Engineer
| Location: | Mumbai |
| Experience: | 3-5 Years |
Role Overview:
We are looking for an experienced L2 Data Engineer responsible for designing, developing, and maintaining data integration and ETL pipelines. The candidate should have hands-on experience in SQL, ETL development, Pentaho Data Integration (PDI), Apache Spark, and Hadoop ecosystem technologies. The role involves working with structured and semi-structured data, optimizing data processing jobs, and supporting enterprise data platforms.
Key Responsibilities:
- Design, develop, test, and maintain ETL/data integration pipelines.
- Develop and optimize complex SQL queries, stored procedures, and database objects.
- Build and maintain data pipelines using Pentaho Data Integration (PDI/Kettle).
- Develop Spark applications for large-scale data processing.
- Work with Hadoop ecosystem components for data storage and processing.
- Perform data extraction, transformation, and loading from multiple source systems.
- Optimize ETL jobs and SQL queries for performance and scalability.
- Perform data quality checks and troubleshoot data issues.
- Monitor production ETL jobs and resolve failures within SLA.
- Collaborate with business analysts, data architects, and application teams to understand data requirements.
- Participate in code reviews and follow development best practices.
- Prepare technical documentation and support deployment activities.
Required Skills and Qualifications:
- Bachelor's degree in Computer Science, Information Technology, or a related field.
- Strong SQL programming skills.
- Hands-on experience with ETL development.
- Experience with Pentaho Data Integration (PDI/Kettle).
- Good knowledge of Apache Spark (Spark SQL, DataFrames).
- Experience working with the Hadoop ecosystem (HDFS, Hive, YARN).
- Understanding of data warehousing concepts and dimensional modeling.
- Experience handling large datasets and optimizing batch processing.
- Familiarity with Linux/Unix commands and shell scripting.
- Experience using Git or other version control systems.
- Strong problem-solving skills and attention to detail.
- Excellent communication and teamwork abilities.
- Ability to work in a fast-paced and dynamic environment.
- Familiarity with Python, Kafka, any cloud platform, file formats (Parquet, ORC, Avro), CI/CD and DevOps practices.