Job Description
Databricks Engineer
Data Lakehouse | Delta Lake | PySpark | MLflow | Unity Catalog
Role Overview
We are looking for a skilled and passionate Databricks Engineer to design,
build, and optimize enterprise-scale data lakehouse solutions on the
Databricks platform. The successful candidate will be responsible for creating
Databricks pipeline delivering Financial Crime platforms covering Anti-Money
Laundering (AML), Know Your Customer (KYC), Customer Risk Assessment (CRA),
Sanctions Screening, Transaction Monitoring, Fraud Detection, and Regulatory
Reporting.
Position Details
Job Title:
Databricks Engineer
Department:
Client Delivery / Data & AI Engineering
Experience Required:
8-10+ Years (overall) | 5+ Years Databricks hands-on
Employment Type:
Full-Time
Location:
Hybrid (as per business requirement)
Reporting To:
Delivery Manager / Technical Project Manager
Key Responsibilities
Databricks Platform Engineering
Design, build, and maintain Databricks workspaces, clusters, and compute
pools across dev/test/prod environments.
Configure and manage Databricks Unity Catalog for data governance, access
control, fine-grained permissions, and data lineage.
Optimize cluster configurations—instance types, auto-scaling policies,
spot/preemptible nodes—for cost and performance.
Implement workspace-level best practices: folder structures, access
controls, secret management (Databricks Secrets / Azure Key Vault / AWS
Secrets Manager).
Manage Databricks jobs, workflows, and multi-task job orchestration with
dependency management.
Delta Lake & Lakehouse Architecture
Design and implement Delta Lake tables with appropriate partitioning,
Z-ordering, and file compaction (OPTIMIZE / VACUUM).
Build Medallion Architecture (Bronze / Silver / Gold) layers for structured
data lake organization.
Implement Delta Live Tables (DLT) pipelines for declarative, reliable
ETL/ELT with built-in data quality expectations.
Manage schema evolution, table versioning, time travel, and Change Data Feed
(CDF) for incremental processing.
Design data lakehouse patterns integrating Delta Lake with external systems
(Kafka, ADLS, S3, GCS).
Data Pipeline Development (PySpark / SQL)
Develop scalable batch and streaming data pipelines using PySpark, Spark
SQL, and Delta Lake.
Build structured streaming pipelines for real-time ingestion from Kafka,
Event Hubs, and Kinesis into Delta tables.
Write optimized PySpark transformations leveraging broadcast joins, adaptive
query execution (AQE), and dynamic partition pruning.
Create reusable transformation libraries, utility frameworks, and pipeline
templates for team productivity.
Implement robust error handling, retry logic, and dead-letter queue patterns
in production pipelines.
MLflow & AI/ML Workloads
Set up and manage MLflow tracking servers, experiment registries, and model
lifecycle management on Databricks.
Support data scientists and ML engineers in deploying model training and
inference workloads on Databricks clusters and GPU instances.
Build feature engineering pipelines using Databricks Feature Store for
reusable, versioned ML features.
Enable GenAI workloads—LLM fine-tuning, RAG pipeline development, and vector
search (Databricks Vector Search / Mosaic AI).
Implement MLOps practices: model versioning, A/B testing, model serving via
Databricks Model Serving endpoints.
Cloud Integration & DevOps
Integrate Databricks with cloud-native services: Azure Data Lake Storage
(ADLS).
Build and maintain CI/CD pipelines for Databricks notebooks and jobs using
Azure DevOps, GitHub Actions, or GitLab CI.
Implement Databricks Asset Bundles (DABs) or Terraform for
infrastructure-as-code (IaC) deployment of Databricks resources.
Manage data ingestion using Auto Loader, COPY INTO, and partner integrations
(Fivetran, dbt, Airbyte).
Monitor pipeline health, cluster utilization, and costs using Databricks
system tables and cloud cost management tools.
Governance, Security & Optimization
Implement row-level security, column masking, and dynamic data views using
Unity Catalog policies.
Ensure data quality enforcement using Delta Live Tables expectations and
Great Expectations integrations.
Conduct performance tuning—query plan analysis, caching strategies, Photon
engine enablement.
Maintain data cataloging, metadata management, and data lineage tracking
within Unity Catalog.
Document architecture decisions, runbooks, and operational guides for
Databricks workloads.
Required Qualifications
Education
Bachelor's or Master's degree in Computer Science, Information Technology,
Data Engineering, or related field.
Experience
8+ years of total experience in data engineering or software engineering.
3+ years of dedicated hands-on experience with the Databricks platform in
production environments.
Strong background in big data engineering, cloud data platforms, and
distributed computing.
Databricks Platform
Deep expertise in Databricks Workspaces, Clusters, Jobs, Workflows, and
Repos.
Proficiency with Unity Catalog—metastore setup, catalog/schema/table
management, access controls, and data lineage.
Hands-on experience with Delta Live Tables (DLT)—pipeline development,
expectations, and monitoring.
Strong command of Delta Lake internals—transaction log, ACID guarantees,
file layout, and optimization techniques.
Experience with Databricks SQL Warehouses, SQL Analytics, and dashboard
creation.
Knowledge of Databricks Photon engine, serverless compute, and cost
optimization strategies.
PySpark & SQL
4+ years of PySpark development—DataFrames, Datasets, Spark SQL, RDD
operations.
Expert-level SQL—window functions, lateral joins, CTEs, recursive queries,
and analytical functions.
Experience with Spark performance tuning—AQE, query plans (EXPLAIN),
partitioning, and caching.
Proficiency with Python for pipeline development, utilities, and automation.
Cloud Platforms
Hands-on experience with at least one: Azure (ADLS Gen2, ADF, Azure
Databricks), AWS (S3, EMR, Glue, AWS Databricks), or GCP (GCS, BigQuery,
Dataproc).
Experience with cloud networking for Databricks: VNet/VPC injection, private
endpoints, and firewall configurations.
Familiarity with IAM roles, managed identities, and service principal
authentication for Databricks.
MLflow & ML Engineering (Nice to Have)
Working knowledge of MLflow—experiment tracking, model registry, and
deployment.
Experience supporting ML pipelines on Databricks for training, evaluation,
and serving.
Exposure to Databricks Feature Store and Mosaic AI / GenAI capabilities.
Preferred Certifications
Databricks Certified Associate Developer for Apache Spark (PySpark or
Scala).
Databricks Certified Data Engineer Associate / Professional—strongly
preferred.
Databricks Certified Machine Learning Associate / Professional.
Azure Data Engineer Associate (DP-203) / AWS Data Analytics Specialty / GCP
Professional Data Engineer.
dbt Analytics Engineer Certification.
Preferred Qualifications
Experience with dbt (data build tool) for SQL-based transformation on
Databricks SQL.
Knowledge of Apache Kafka / Confluent for real-time streaming into
Databricks.
Familiarity with Terraform or Pulumi for Databricks infrastructure-as-code.
Exposure to Apache Iceberg or Apache Hudi in addition to Delta Lake.
Experience with BI tool integration: Power BI, Tableau, or Looker connected
to Databricks SQL.
Knowledge of data mesh principles and federated data governance at scale.
Domain experience in BFSI, Healthcare, Retail, or Manufacturing data
programs.
Core Competencies
Deep technical depth in Databricks with the ability to architect and
troubleshoot complex lakehouse systems.
Strong problem-solving skills—ability to diagnose pipeline failures,
performance bottlenecks, and data quality issues.
Collaborative team player comfortable working with data engineers, ML
engineers, and business stakeholders.
Excellent documentation and communication skills for technical and
non-technical audiences.
Continuous learner—actively keeps pace with Databricks platform updates and
the broader data/AI ecosystem.
Delivery-oriented mindset with experience in Agile/Scrum delivery
environments.
What We Offer
Work on large-scale, production Databricks implementations for marquee
enterprise clients.
Exposure to the full Databricks ecosystem—lakehouse, streaming, GenAI, and
MLOps.
Databricks certification sponsorship and continuous learning support.
Competitive compensation, performance bonuses, and comprehensive benefits.
Hybrid/remote work flexibility and inclusive, high-performance team culture.
Direct collaboration with Databricks account teams and technical partners.