Databricks Engineer
Experience:
4–6 Years (Minimum 2+ Years of Hands-on Databricks Experience)
Location:
Hybrid
Employment Type:
Full-Time
About the Role
We are seeking a highly skilled
Databricks Engineer
to design, build, and optimize enterprise-scale Lakehouse solutions on the
Databricks platform. The ideal candidate will have strong expertise in
Databricks, PySpark, SQL, Delta Lake, and cloud data platforms
, along with experience in building scalable batch and streaming data
pipelines.
In this role, you will contribute to the development of mission-critical
Financial Crime platforms supporting
AML, KYC, Customer Risk Assessment, Sanctions Screening, Transaction
Monitoring, Fraud Detection, and Regulatory Reporting
while leveraging modern Lakehouse architecture and Data Engineering best
practices.
Key Responsibilities
Databricks Platform Engineering
-
Design, develop, and maintain Databricks workspaces, clusters, jobs,
workflows, and compute environments.
-
Configure and manage
Unity Catalog
for centralized governance, security, metadata management, and data
lineage.
-
Optimize cluster configurations for performance, scalability, and cost
efficiency.
-
Implement secure access management, secrets management, and workspace
governance.
Data Engineering & Lakehouse Architecture
-
Design and implement scalable
Delta Lake
solutions using Medallion Architecture (Bronze, Silver, Gold).
-
Develop batch and real-time data pipelines using
PySpark, Spark SQL, and Delta Lake
.
-
Build and maintain
Delta Live Tables (DLT)
pipelines with automated data quality checks.
-
Implement schema evolution, Change Data Feed (CDF), partitioning,
optimization, and time travel capabilities.
-
Develop streaming pipelines using Kafka, Event Hubs, or similar
technologies.
Performance Optimization
-
Optimize Spark workloads using partitioning, caching, Adaptive Query
Execution (AQE), broadcast joins, and Photon Engine.
-
Monitor and improve pipeline reliability, scalability, and operational
efficiency.
-
Implement robust error handling, retry mechanisms, and production
monitoring.
AI/ML & MLflow (Preferred)
-
Support ML workloads using
MLflow
for experiment tracking and model lifecycle management.
-
Enable feature engineering pipelines and collaborate with ML teams.
-
Exposure to GenAI technologies, RAG pipelines, Vector Search, or Mosaic AI
is an added advantage.
Cloud & DevOps
-
Work with Azure, AWS, or GCP cloud services for enterprise data
engineering solutions.
-
Build CI/CD pipelines using Azure DevOps, GitHub Actions, or GitLab CI.
-
Implement Infrastructure as Code (Terraform/Databricks Asset Bundles)
where applicable.
-
Develop automated ingestion pipelines using Auto Loader, COPY INTO, dbt,
or similar tools.
Governance & Security
-
Implement data governance using Unity Catalog.
-
Configure row-level security, column masking, and fine-grained access
controls.
-
Maintain metadata, lineage, and documentation for enterprise data assets.
Required Skills
-
Databricks
-
Python
-
PySpark
-
SQL
-
Delta Lake
-
Unity Catalog
-
Delta Live Tables (DLT)
-
Spark SQL
-
Lakehouse Architecture
-
Data Engineering
-
Performance Tuning
-
ETL/ELT Pipelines
-
Git
Cloud Technologies
Experience with at least one of the following:
-
Azure:
ADLS Gen2, Azure Databricks, Azure Data Factory
-
AWS:
S3, EMR, Glue, AWS Databricks
-
GCP:
GCS, BigQuery, Dataproc
Good to Have
-
MLflow
-
MLOps
-
AI / GenAI / LLM Exposure
-
Databricks Feature Store
-
Kafka / Event Hubs
-
Terraform
-
dbt
-
Auto Loader
-
Power BI / Tableau / Looker
-
Apache Iceberg or Apache Hudi
Qualifications
-
Bachelor's or Master's degree in Computer Science, Information Technology,
Data Engineering, or a related discipline.
-
4–6 years of experience in Data Engineering or Software Engineering.
-
Minimum 2 years of hands-on production experience with Databricks.
Preferred Certifications
-
Databricks Certified Data Engineer Associate / Professional
-
Databricks Certified Associate Developer for Apache Spark
-
Azure Data Engineer Associate (DP-203)
-
AWS Data Analytics Specialty
-
GCP Professional Data Engineer
Why Join Us?
-
Work on enterprise-scale Databricks and Lakehouse implementations.
-
Exposure to modern Data Engineering, Streaming, AI/ML, and MLOps
technologies.
-
Opportunity to work on high-impact Financial Crime and Analytics
platforms.
-
Continuous learning, certification support, and career growth.
-
Competitive compensation and a collaborative, high-performance work
environment.