POSITION SUMMARY:
We are seeking a Data Engineer with 5+ years of experience in building and
maintaining data pipelines using PySpark, SQL, and Python. The candidate
should have a solid understanding of Bigdata tools like Hadoop, Hive, Oozie
and have basic knowledge of cloud environments (preferably AWS). The role
requires working closely with teams to support data processing and analytics
needs
ROLES AND RESPONSIBILITIES:
• Develop and maintain data pipelines using PySpark and Python, Hive, Oozie
• Write efficient SQL queries for data extraction, transformation, and validation
• Assist in integrating data from multiple sources and ensuring data accuracy
• Support debugging, monitoring, and optimization of data pipelines
• Collaborate with team members to understand data requirements and deliver solutions
• Follow best practices for data engineering and doc
REQUIRED SKILLS:
• 5+ years of hands-on experience in PySpark, SQL, and Python
• 5+ years of hands-on experience in Big Data Tools like Hive ,Oozie
• Working knowledge of cloud environments (preferably AWS) • Understanding
of data processing and ETL concept
SECONDARY SKILLS:
• Basic experience with Databricks
• Familiarity with Linux/Unix commands for working on edge nodes
• Exposure to scheduling or orchestration tools
Apply through whichever channel suits you best.