Job Description – Data Engineer (PySpark & Azure)
Position:
Data Engineer
Experience:
6 - 8 Years
Location:
Pune / Bangalore
Work Mode:
Hybrid
Notice Period:
Immediate to 30 Days Preferred
Role Overview
We are looking for a skilled
Data Engineer
with hands-on experience in
PySpark
and
Microsoft Azure
to design, develop, and optimize scalable data pipelines. The ideal
candidate should have expertise in big data technologies, cloud-based data
engineering, and ETL development to support enterprise analytics and
data-driven decision-making.
Key Responsibilities
-
Design, build, and maintain scalable ETL/ELT pipelines using
PySpark
.
-
Develop and optimize data processing workflows on the
Microsoft Azure
platform.
-
Build and manage data ingestion pipelines from multiple structured and
unstructured data sources.
-
Work with Azure services such as
Azure Data Factory (ADF), Azure Databricks, Azure Data Lake Storage
(ADLS), and Azure Synapse Analytics
.
-
Develop efficient data models and ensure high data quality, integrity,
and governance.
-
Optimize PySpark jobs for performance, scalability, and cost efficiency.
-
Collaborate with data analysts, data scientists, and business
stakeholders to understand data requirements.
-
Implement data validation, monitoring, and troubleshooting processes.
-
Ensure adherence to data security, governance, and compliance standards.
-
Participate in code reviews and contribute to best practices for data
engineering.
Required Skills & Qualifications
-
Bachelor's or Master's degree in Computer Science, Information
Technology, Engineering, or a related field.
-
6 - 8 years
of experience as a Data Engineer.
-
Strong hands-on experience with
PySpark
(Mandatory).
-
Experience working on the
Microsoft Azure
cloud platform (Mandatory).
-
Expertise in
Azure Databricks
,
Azure Data Factory (ADF)
,
Azure Data Lake Storage (ADLS)
, and
Azure Synapse Analytics
.
-
Strong SQL skills and experience with relational databases.
-
Experience with Python for data engineering and automation.
-
Knowledge of data warehousing concepts and ETL/ELT methodologies.
-
Familiarity with Git, Azure DevOps, or CI/CD pipelines.
-
Strong analytical and problem-solving skills.
Preferred Skills
-
Experience with Delta Lake and Apache Spark optimization.
-
Exposure to Snowflake, Kafka, or other cloud data platforms.
-
Knowledge of Power BI or other reporting tools.
-
Experience working in Agile/Scrum environments.