Databricks Platform Manager / Engineer (AWS)
| Experience | 5 – 10 years of overall IT experience, including 3 – 5+ years of hands-on experience administering and managing Databricks platforms in production environments on AWS. |
Job Summary
We are looking for a Databricks Platform Manager/Engineer responsible for designing, administering, securing, and optimizing the enterprise Databricks platform on AWS. The ideal candidate will own the platform lifecycle, ensuring scalability, security, governance, cost optimization, and operational excellence while enabling data engineering, analytics, and AI teams to work efficiently. This role focuses on platform engineering rather than data engineering or analytics development.
Key Responsibilities:
Platform Administration
- Design, deploy, and manage Databricks workspaces across Development, Test, UAT, and Production environments.
- Establish workspace standards, networking, and environment isolation.
- Manage platform upgrades, including Databricks Runtime (DBR) version planning, testing, and rollout.
- Maintain platform availability, reliability, and performance.
Compute & Cluster Management
- Create and manage cluster policies to enforce organizational standards.
- Configure instance pools and job clusters.
- Prevent oversized or non-compliant clusters.
- Standardize cluster configurations across teams.
- Optimize cluster startup times and resource utilization.
Security & Identity Management
- Integrate Databricks with AWS IAM.
- Configure and manage authentication and authorization.
- Manage service principals, personal access tokens (PATs), and secret scopes.
- Implement least-privilege access controls.
- Support enterprise identity management and SSO.
Governance & Data Access
- Administer Unity Catalog.
- Manage catalogs, schemas, tables, and permissions.
- Implement role-based access control (RBAC).
- Enforce data governance and access policies.
- Support audit and compliance requirements.
DevOps & CI/CD
- Configure Git integration with Databricks Repos.
- Support CI/CD pipelines for notebooks, workflows, and infrastructure.
- Collaborate with DevOps teams on automated deployments.
- Enable Infrastructure as Code (IaC) using Terraform where applicable.
Monitoring & Operations:
- Monitor workspace health, jobs, clusters, and platform performance.
- Respond to incidents and perform root cause analysis.
- Maintain operational dashboards and alerts.
- Support production releases and platform maintenance activities.
Cost Optimization:
- Monitor Databricks usage and cloud spending.
- Recommend cluster sizing and auto-scaling strategies.
- Optimize job scheduling and compute utilization.
- Track cost trends and recommend savings opportunities.
Compliance & Audit:
- Configure audit logging.
- Ensure platform complies with enterprise security policies.
- Support internal and external audits.
- Maintain operational documentation and platform standards.
Required Skills:
- Databricks Platform
- Workspace administration
- Cluster management
- Cluster policies
- Instance pools
- Job clusters
- Databricks Runtime management
- Workflows
- Databricks Repos
- Unity Catalog
- Secret Scopes
- Service Principals
- Personal Access Tokens (PATs)
- AWS
- IAM
- VPC fundamentals
- EC2 concepts
- S3
- CloudWatch
- KMS
- Security Groups
- Networking fundamentals
- DevOps
- Git
- CI/CD
- Terraform
- Infrastructure as Code
- Azure DevOps, GitHub Actions, or Jenkins
- Monitoring
- Platform monitoring
- Incident management
- Log analysis
- Performance tuning
- Cost optimization
Required Experience:
Candidates should have hands-on experience with:
- Managing enterprise Databricks platforms
- Setting up multiple Databricks workspaces
- Configuring Unity Catalog
- AWS IAM integration
- Cluster policy creation and governance
- Runtime version upgrades
- CI/CD implementation
- Git integration
- Platform monitoring and troubleshooting
- Secrets management
- Audit logging
- Production support
- Cost optimization initiatives
Preferred Qualifications:
- Experience managing large-scale Databricks environments (100+ users).
- Experience supporting multiple business units or enterprise data platforms.
- Hands-on experience with Terraform for Databricks infrastructure provisioning.
- Knowledge of Lakehouse architecture and Delta Lake.
- Experience with FinOps practices for cloud cost optimization.
- Familiarity with data governance and security frameworks.
Nice to Have:
- Experience with AWS Organizations and Control Tower.
- Knowledge of networking concepts for private connectivity (PrivateLink, VPC endpoints).
- Experience with monitoring tools such as CloudWatch, Datadog, or Splunk.
- Exposure to Apache Spark internals.
- Experience supporting AI/ML workloads on Databricks.
Education
Bachelor's degree in Computer Science, Information Technology, Engineering, or a related field.
Certifications (Preferred):
- Databricks Certified Data Engineer Professional or Databricks Certified Platform Administrator (if available).
- AWS Certified Solutions Architect – Associate/Professional.
- AWS Certified SysOps Administrator.
- Terraform Associate certification.
Ideal Candidate Profile:
The ideal candidate has a strong platform engineering and cloud infrastructure background with deep expertise in Databricks administration on AWS. They are comfortable owning the operational health, security, governance, automation, and cost optimization of the Databricks platform while partnering closely with data engineering, analytics, DevOps, and security teams. Experience with Terraform is highly desirable, as many enterprise Databricks environments are provisioned and managed through Infrastructure as Code.
Document generated for internal use only. Please refer to the official documentation for comprehensive guidelines and policies.