JOB DESCRIPTION
Databricks Engineer
Data Lakehouse
|
Delta Lake
|
PySpark
|
MLflow
|
Unity Catalog
Role Overview
We are looking for a skilled and passionate Databricks Engineer to design,
build, and optimize enterprise-scale data lakehouse solutions on the
Databricks platform. The successful candidate will be responsible for
creating Databricks pipeline delivering Financial Crime platforms covering
Anti-Money Laundering (AML), Know Your Customer (KYC), Customer Risk
Assessment (CRA), Sanctions Screening, Transaction Monitoring, Fraud
Detection, and Regulatory Reporting
Position Details
Job Title:
Databricks Engineer
Department:
Client Delivery / Data & AI Engineering
Experience Required:
8-10
+ Years (overall) | 5+ Years Databricks hands-on
Employment Type:
Full-Time
Location:
Hybrid (as per business requirement)
Reporting To:
Delivery Manager / Technical Project Manager
Key Responsibilities
Databricks Platform Engineering
●
Design, build, and maintain Databricks workspaces, clusters, and compute
pools across dev/test/prod environments.
●
Configure and manage Databricks Unity Catalog for data governance, access
control, fine-grained permissions, and data lineage.
●
Optimize cluster configurations — instance types, auto-scaling policies,
spot/preemptible nodes — for cost and performance.
●
Implement workspace-level best practices: folder structures, access
controls, secret management (Databricks Secrets / Azure Key Vault / AWS
Secrets Manager).
●
Manage Databricks jobs, workflows, and multi-task job orchestration with
dependency management.
Delta Lake & Lakehouse Architecture
●
Design and implement Delta Lake tables with appropriate partitioning,
Z-ordering, and file compaction (OPTIMIZE / VACUUM).
●
Build Medallion Architecture (Bronze / Silver / Gold) layers for
structured data lake organization.
●
Implement Delta Live Tables (DLT) pipelines for declarative, reliable
ETL/ELT with built-in data quality expectations.
●
Manage schema evolution, table versioning, time travel, and Change Data
Feed (CDF) for incremental processing.
●
Design data lakehouse patterns integrating Delta Lake with external
systems (Kafka, ADLS, S3, GCS).
Data Pipeline Development (PySpark / SQL)
●
Develop scalable batch and streaming data pipelines using PySpark, Spark
SQL, and Delta Lake.
●
Build structured streaming pipelines for real-time ingestion from Kafka,
Event Hubs, and Kinesis into Delta tables.
●
Write optimized PySpark transformations leveraging broadcast joins,
adaptive query execution (AQE), and dynamic partition pruning.
●
Create reusable transformation libraries, utility frameworks, and pipeline
templates for team productivity.
●
Implement robust error handling, retry logic, and dead-letter queue
patterns in production pipelines.
MLflow & AI/ML Workloads
●
Set up and manage MLflow tracking servers, experiment registries, and
model lifecycle management on Databricks.
●
Support data scientists and ML engineers in deploying model training and
inference workloads on Databricks clusters and GPU instances.
●
Build feature engineering pipelines using Databricks Feature Store for
reusable, versioned ML features.
●
Enable GenAI workloads — LLM fine-tuning, RAG pipeline development, and
vector search (Databricks Vector Search / Mosaic AI).
●
Implement MLOps practices: model versioning, A/B testing, model serving
via Databricks Model Serving endpoints.
Cloud Integration & DevOps
●
Integrate Databricks with cloud-native services: Azure Data Lake Storage
(ADLS).
●
Build and maintain CI/CD pipelines for Databricks notebooks and jobs using
Azure DevOps, GitHub Actions, or GitLab CI.
●
Implement Databricks Asset Bundles (DABs) or Terraform for
infrastructure-as-code (IaC) deployment of Databricks resources.
●
Manage data ingestion using Auto Loader, COPY INTO, and partner
integrations (Fivetran, dbt, Airbyte).
●
Monitor pipeline health, cluster utilization, and costs using Databricks
system tables and cloud cost management tools.
Governance, Security & Optimization
●
Implement row-level security, column masking, and dynamic data views using
Unity Catalog policies.
●
Ensure data quality enforcement using Delta Live Tables expectations and
Great Expectations integrations.
●
Conduct performance tuning — query plan analysis, caching strategies,
Photon engine enablement.
●
Maintain data cataloging, metadata management, and data lineage tracking
within Unity Catalog.
●
Document architecture decisions, runbooks, and operational guides for
Databricks workloads.
Required Qualifications
Education
●
Bachelor's or Master's degree in Computer Science, Information Technology,
Data Engineering, or related field.
Experience
●
8+ years of total experience in data engineering or software engineering.
●
3+ years of dedicated hands-on experience with the Databricks platform in
production environments.
●
Strong background in big data engineering, cloud data platforms, and
distributed computing.
Databricks Platform
●
Deep expertise in Databricks Workspaces, Clusters, Jobs, Workflows, and
Repos.
●
Proficiency with Unity Catalog — metastore setup, catalog/schema/table
management, access controls, and data lineage.
●
Hands-on experience with Delta Live Tables (DLT) — pipeline development,
expectations, and monitoring.
●
Strong command of Delta Lake internals — transaction log, ACID guarantees,
file layout, and optimization techniques.
●
Experience with Databricks SQL Warehouses, SQL Analytics, and dashboard
creation.
●
Knowledge of Databricks Photon engine, serverless compute, and cost
optimization strategies.
PySpark & SQL
●
4+ years of PySpark development — DataFrames, Datasets, Spark SQL, RDD
operations.
●
Expert-level SQL — window functions, lateral joins, CTEs, recursive
queries, and analytical functions.
●
Experience with Spark performance tuning — AQE, query plans (EXPLAIN),
partitioning, and caching.
●
Proficiency with Python for pipeline development, utilities, and
automation.
Cloud Platforms
●
Hands-on experience with at least one: Azure (ADLS Gen2, ADF, Azure
Databricks), AWS (S3, EMR, Glue, AWS Databricks), or GCP (GCS, BigQuery,
Dataproc).
●
Experience with cloud networking for Databricks: VNet/VPC injection,
private endpoints, and firewall configurations.
●
Familiarity with IAM roles, managed identities, and service principal
authentication for Databricks.
MLflow & ML Engineering (Nice to Have)
●
Working knowledge of MLflow — experiment tracking, model registry, and
deployment.
●
Experience supporting ML pipelines on Databricks for training, evaluation,
and serving.
●
Exposure to Databricks Feature Store and Mosaic AI / GenAI capabilities.
Preferred Certifications
●
Databricks Certified Associate Developer for Apache Spark (PySpark or
Scala).
●
Databricks Certified Data Engineer Associate / Professional — strongly
preferred.
●
Databricks Certified Machine Learning Associate / Professional.
●
Azure Data Engineer Associate (DP-203) / AWS Data Analytics Specialty /
GCP Professional Data Engineer.
●
dbt Analytics Engineer Certification.
Preferred Qualifications
●
Experience with dbt (data build tool) for SQL-based transformation on
Databricks SQL.
●
Knowledge of Apache Kafka / Confluent for real-time streaming into
Databricks.
●
Familiarity with Terraform or Pulumi for Databricks
infrastructure-as-code.
●
Exposure to Apache Iceberg or Apache Hudi in addition to Delta Lake.
●
Experience with BI tool integration: Power BI, Tableau, or Looker
connected to Databricks SQL.
●
Knowledge of data mesh principles and federated data governance at scale.
●
Domain experience in BFSI, Healthcare, Retail, or Manufacturing data
programs.
Core Competencies
●
Deep technical depth in Databricks with the ability to architect and
troubleshoot complex lakehouse systems.
●
Strong problem-solving skills — ability to diagnose pipeline failures,
performance bottlenecks, and data quality issues.
●
Collaborative team player comfortable working with data engineers, ML
engineers, and business stakeholders.
●
Excellent documentation and communication skills for technical and
non-technical audiences.
●
Continuous learner — actively keeps pace with Databricks platform updates
and the broader data/AI ecosystem.
●
Delivery-oriented mindset with experience in Agile/Scrum delivery
environments.
What We Offer
●
Work on large-scale, production Databricks implementations for marquee
enterprise clients.
●
Exposure to the full Databricks ecosystem — lakehouse, streaming, GenAI,
and MLOps.
●
Databricks certification sponsorship and continuous learning support.
●
Competitive compensation, performance bonuses, and comprehensive benefits.
●
Hybrid/remote work flexibility and inclusive, high-performance team
culture.
●
Direct collaboration with Databricks account teams and technical partners.
Confidential
|
Client Delivery
|
EXL Service
|
2026
Apply through whichever channel suits you best.