Role Overview
We are seeking an ambitious Data and AI Engineer with 5-8 years of experience
to
play a pivotal role in modernizing client’s core data infrastructure and
scaling advanced
AI capabilities. This role bridges legacy big data environment and future
cloud platform.
You will actively maintain and optimize our existing data pipelines while
architecting,
developing, and transitioning workflows to an AI-ready cloud ecosystem. Beyond
traditional data pipelines, you will play an active role in building and
deploying intelligent
AI agents and leveraging advanced Large Language Models (LLMs) like Anthropic
Claude, utilizing the cutting-edge Databricks AI suite to deliver immediate
business
value.
Experience Level: 5-8 Years
Current Architecture: Hadoop (HDFS), Apache Spark, MapReduce, Hive
Target Architecture: AWS (S3, Glue, EMR), Hive, HDFS, Airflow, ControlM,
Databricks (Delta Lake, Unity Catalog, Vector Search, Mosaic AI)
Core AI Stack: Anthropic Claude, Copilot, LangChain GenAI Agents
BI & Analytics Stack: Tableau Desktop & Tableau Server
Core Responsibilities
Cloud Modernization & Migration: Deconstruct legacy Apache Spark and
Hadoop MapReduce workflows to re-architect and rebuild them as optimized,
production-ready pipelines within AWS and Databricks.
AI Agent & LLM Development: Design, build, and deploy intelligent AI
agents
and workflow automation tools leveraging leading Large Language Models
(LLMs) such as Anthropic Claude.
AI Data Pipeline Engineering: Build and optimize pipeline architectures
explicitly tailored for AI use cases, including unstructured data ingestion,
real-
time feature tokenization, and metadata tagging for vector databases.
Legacy Infrastructure Maintenance: Monitor, maintain, and troubleshoot
existing big data workloads running on our Hadoop cluster to guarantee data
availability for business operations during the multi-phase migration,
resolving
bottlenecks and Out-Of-Memory (OOM) errors.
BI Engineering & Support: Act as the primary engineering liaison for
downstream business stakeholders utilizing BI tools (e.g. Tableau, Looker,
etc.)
by performing minor functional enhancements, bug fixes, and data extract
optimizations to resolve report dashboard latency.
Cloud Optimization: Utilize Databricks and Delta Lake features (e.g., ACID
transactions, Z-Ordering, caching) to significantly improve pipeline
performance,
reliability, and cost efficiency.
Databricks AI Suite Implementation: Leverage Databricks tools (such as
Databricks Vector Search, Mosaic AI, and Lakeflow) to orchestrate, track, and
serve production-grade Generative AI and Retrieval-Augmented Generation
(RAG) applications.
Required Qualifications
Experience: 5 to 8 years of professional software engineering or data
engineering experience in a production environment.
AI & Agentic Frameworks: Hands-on experience or deep technical
familiarity
building functional AI agents, integrating LLM APIs (specifically Anthropic
Claude), and utilizing orchestration frameworks (e.g., LangChain, Databricks
Mosaic AI Agent Framework or any other tool ).
Distributed Computing: Foundational understanding of distributed storage and
computing concepts—specifically partitioning, shuffling, caching, and
broadcast
joins. Solid hands-on experience with Apache Spark is required.
Programming & SQL: Strong proficiency in Python (PySpark) or Scala,
alongside intermediate-to-advanced SQL querying capabilities (window
functions, query tuning, and complex joins).
Cloud & Databricks Exposure: Direct experience or deep theoretical
knowledge of the AWS ecosystem (S3, IAM) and Databricks environments.
Visualization Layer: Practical experience working with any BI tool (e.g.
Tableau,
Looker, Power BI, etc.) with the capability to debug calculated fields, modify
parameters, and troubleshoot slow-loading reports.
Education: Bachelor’s degree in Computer Science, Data Engineering,
Information Systems, or a related quantitative field, or equivalent practical
experience.
Apply through whichever channel suits you best.