Job Description – Lead Big Data Engineer
Job Title:
Lead Big Data Engineer
Experience:
7–9 Years
Location:
All EXL Locations (Hybrid)
About the Role
We are looking for a
Lead Big Data Engineer
with 7–9 years of experience in designing, developing, and optimizing
large-scale data processing solutions. The ideal candidate will have strong
expertise in
Python, Apache Spark, Hadoop ecosystem, Hive, and SQL
, with experience building scalable data pipelines and processing high-volume
structured and unstructured data. The role requires close collaboration with
cross-functional teams to deliver robust, high-performance data engineering
solutions.
Key Responsibilities
-
Design, build, and maintain scalable big data pipelines using
Python and Apache Spark
.
-
Develop and optimize batch and distributed data processing applications on
the
Hadoop ecosystem
.
-
Build and manage data warehousing solutions using
Hive
and write efficient SQL queries for data analysis and transformation.
-
Optimize Spark jobs for performance, scalability, and resource utilization.
-
Collaborate with data analysts, data scientists, and business stakeholders
to understand data requirements and deliver reliable solutions.
-
Perform data ingestion, cleansing, transformation, and validation across
multiple data sources.
-
Troubleshoot production issues and ensure high availability and reliability
of data platforms.
-
Lead code reviews, mentor junior engineers, and drive engineering best
practices.
-
Ensure data quality, governance, security, and compliance standards are
followed.
-
Participate in Agile ceremonies and contribute to solution design and
technical discussions.
Required Skills
-
7–9 years of experience in Big Data Engineering.
-
Strong programming experience in
Python
.
-
Hands-on expertise with
Apache Spark
for large-scale data processing.
-
Experience working with the
Hadoop ecosystem
(HDFS, YARN, MapReduce).
-
Strong knowledge of
Hive
for data warehousing and querying.
-
Advanced SQL skills for data manipulation, optimization, and performance
tuning.
-
Experience building ETL/ELT pipelines for large datasets.
-
Strong understanding of distributed computing concepts and data processing
frameworks.
-
Experience with version control systems such as Git.
-
Excellent analytical, problem-solving, and debugging skills.
Preferred Skills
-
Experience with workflow orchestration tools such as Airflow or Oozie.
-
Exposure to cloud platforms (AWS, Azure, or GCP).
-
Knowledge of Kafka or other streaming technologies.
-
Familiarity with CI/CD practices and DevOps concepts.
-
Experience working in Agile/Scrum environments.
Education
-
Bachelor's or Master's degree in Computer Science, Information Technology,
Engineering, or a related field.