Position Summary
We are seeking an experienced
Lead Data Engineer
with
7–10 years of hands-on experience
in designing, developing, and maintaining scalable data pipelines and
processing systems within an
on-premises Big Data environment
.
This role combines strong technical expertise with leadership
responsibilities. The ideal candidate will own architecture and design
decisions, drive technical excellence, mentor team members, and collaborate
with cross-functional stakeholders to deliver reliable, high-quality data
solutions that support analytics and business needs.
Key Responsibilities
-
Lead the design, development, and maintenance of scalable data pipelines
for large-scale data ingestion, transformation, and processing.
-
Own architecture and design decisions, define technical standards, and
evaluate technology trade-offs.
-
Mentor and support data engineers through code reviews, design reviews,
technical guidance, and hands-on problem-solving.
-
Develop and optimize data workflows using
Python, Apache Spark, and Hadoop ecosystem technologies
.
-
Work extensively with
Hive, HDFS, and Impala
to manage and analyze large datasets.
-
Manage batch scheduling and workflow orchestration using enterprise
schedulers such as
CA7
or
Control-M
.
-
Ensure data quality, integrity, scalability, and platform performance.
-
Partner with data analysts, data scientists, and business stakeholders to
translate business requirements into technical solutions.
-
Troubleshoot complex production issues and serve as the technical
escalation point for the team.
-
Promote best practices in coding standards, testing, version control,
documentation, and performance optimization.
-
Stay updated on industry trends and emerging technologies, including
AI/ML
, and identify opportunities to enhance data engineering workflows.
Required Qualifications
-
7–10 years of experience in Data Engineering with demonstrated technical
leadership and design ownership.
-
Strong hands-on expertise in
Python
for building enterprise-grade data solutions.
-
Extensive experience with the
Hadoop ecosystem
, including:
-
Hadoop
-
Hive
-
HDFS
-
Impala
-
Strong experience developing and optimizing distributed data processing
applications using
Apache Spark
.
-
Hands-on experience with job scheduling and workflow orchestration tools
such as
CA7
,
Control-M
, or similar enterprise schedulers.
-
Strong knowledge of:
-
SQL
-
ETL/ELT processes
-
Data structures
-
Large-scale data processing frameworks
-
Proven ability to mentor developers, review designs, and make sound
architectural decisions.
-
Exposure to AI/ML concepts, tools, or frameworks with a willingness to
further expand expertise in this area.