Databricks Engineer
EXL Talent Acquisition Team- Location
- Chennai, Tamil Nadu, India
- Workplace
- Hybrid
- Employment
- Full Time
- Salary
- —
Posted 1mo ago
We are looking for a skilled and passionate Databricks Engineer to design, build, and optimize enterprise-scale data lakehouse solutions on the Databricks platform. The successful candidate will be responsible for creating Databricks pipeline delivering Financial Crime platforms covering Anti-Money Laundering (AML), Know Your Customer (KYC), Customer Risk Assessment (CRA), Sanctions Screening, Transaction Monitoring, Fraud Detection, and Regulatory Reporting
Databricks Platform Engineering
- Design, build, and maintain Databricks workspaces, clusters, and compute pools across dev/test/prod environments.
- Configure and manage Databricks Unity Catalog for data governance, access control, fine-grained permissions, and data lineage.
- Optimize cluster configurations — instance types, auto-scaling policies, spot/preemptible nodes — for cost and performance.
- Implement workspace-level best practices: folder structures, access controls, secret management (Databricks Secrets / Azure Key Vault / AWS Secrets Manager).
- Manage Databricks jobs, workflows, and multi-task job orchestration with dependency management.
Delta Lake & Lakehouse Architecture
- Design and implement Delta Lake tables with appropriate partitioning, Z-ordering, and file compaction (OPTIMIZE / VACUUM).
- Build Medallion Architecture (Bronze / Silver / Gold) layers for structured data lake organization.
- Implement Delta Live Tables (DLT) pipelines for declarative, reliable ETL/ELT with built-in data quality expectations.
- Manage schema evolution, table versioning, time travel, and Change Data Feed (CDF) for incremental processing.
- Design data lakehouse patterns integrating Delta Lake with external systems (Kafka, ADLS, S3, GCS).
Data Pipeline Development (PySpark / SQL)
- Develop scalable batch and streaming data pipelines using PySpark, Spark SQL, and Delta Lake.
- Build structured streaming pipelines for real-time ingestion from Kafka, Event Hubs, and Kinesis into Delta tables.
- Write optimized PySpark transformations leveraging broadcast joins, adaptive query execution (AQE), and dynamic partition pruning.
- Create reusable transformation libraries, utility frameworks, and pipeline templates for team productivity.
- Implement robust error handling, retry logic, and dead-letter queue patterns in production pipelines.
MLflow & AI/ML Workloads
- Set up and manage MLflow tracking servers, experiment registries, and model lifecycle management on Databricks.
- Support data scientists and ML engineers in deploying model training and inference workloads on Databricks clusters and GPU instances.
- Build feature engineering pipelines using Databricks Feature Store for reusable, versioned ML features.
- Enable GenAI workloads — LLM fine-tuning, RAG pipeline development, and vector search (Databricks Vector Search / Mosaic AI).
- Implement MLOps practices: model versioning, A/B testing, model serving via Databricks Model Serving endpoints.
Cloud Integration & DevOps
- Integrate Databricks with cloud-native services: Azure Data Lake Storage (ADLS).
- Build and maintain CI/CD pipelines for Databricks notebooks and jobs using Azure DevOps, GitHub Actions, or GitLab CI.
- Implement Databricks Asset Bundles (DABs) or Terraform for infrastructure-as-code (IaC) deployment of Databricks resources.
- Manage data ingestion using Auto Loader, COPY INTO, and partner integrations (Fivetran, dbt, Airbyte).
- Monitor pipeline health, cluster utilization, and costs using Databricks system tables and cloud cost management tools.
Governance, Security & Optimization
- Implement row-level security, column masking, and dynamic data views using Unity Catalog policies.
- Ensure data quality enforcement using Delta Live Tables expectations and Great Expectations integrations.
- Conduct performance tuning — query plan analysis, caching strategies, Photon engine enablement.
- Maintain data cataloging, metadata management, and data lineage tracking within Unity Catalog.
- Document architecture decisions, runbooks, and operational guides for Databricks workloads.
Education
- Bachelor's or Master's degree in Computer Science, Information Technology, Data Engineering, or related field.
Experience
- 4-6 years of total experience in data engineering or software engineering.
- 2+ years of dedicated hands-on experience with the Databricks platform in production environments.
- Strong background in big data engineering, cloud data platforms, and distributed computing.
Skills
- Databricks
- Unity Catalog
- Azure Key Vault
- Secrets Manager
- Delta Lake
- Delta Live Tables
- ETL
- ELT
- Kafka
- S3
- Google Cloud Storage
- PySpark
- SQL
- Spark
- Azure Event Hubs
- AWS Kinesis
- MLflow
- Machine Learning
- Generative AI
- LLM
- Retrieval-Augmented Generation
- MLOps
- Azure Data Lake Storage
- Azure DevOps
- GitHub Actions
- GitLab CI
- Terraform
- Fivetran
- dbt
- Airbyte
- Great Expectations
More jobs at EXL Talent Acquisition Team
All 336Senior Test Automation Engineer
EXL Talent Acquisition Team · Ciudad de México, Mexico, Mexico · today
Senior DevOps Engineer
EXL Talent Acquisition Team · Ciudad de México, Mexico, Mexico · today
Sr Data Engineer
EXL Talent Acquisition Team · Gurugram, Haryana, India · today
Assistant Manager
EXL Talent Acquisition Team · Hyderabad, Telangana, India · today
Data Analyst
EXL Talent Acquisition Team · London, England, United Kingdom · today
Similar roles
Analyst IT – Quality
Mattel · Hyderabad, India · today
Principal Network Engineer
Hewlett Packard Enterprise · Bangalore, Karnataka, India · today
Principal Network Engineer
Hpe · Bangalore, Karnataka, India · today
Cloud DevOps Engineer - Data QA
Southwest Airlines · India Office · today
DevOps Engineer
Ensono · Bengaluru, India · Chennai, India · Hyderabad, India +1 · today
DevOps Engineer
Sonicwall · Bengaluru, Karnataka, India · today