JobHabor

Senior Data Engineer - Databricks & Streaming - Healthcare AI (Onsite, Evening Shift, Lahore, PKR Salary)

HR POD - Hiring Talent Globally
Location
Lahore, Pakistan
Workplace
Employment
Full Time
Salary
Apply on the employer’s site

Posted yesterday

Requirements

  • 4+ years of experience in data engineering, with substantial production experience in Databricks.
  • Strong experience with Spark SQL, PySpark, Delta Lake, Medallion Architecture, and Delta Live Tables (DLT).
  • Hands-on experience with Structured Streaming or equivalent production-grade streaming ingestion using Azure Event Hubs, Kafka, or Kinesis.
  • Strong understanding of checkpoint recovery, watermarking, and deduplication strategies.
  • Demonstrated experience debugging source-to-warehouse data discrepancies.
  • Ability to walk through a real-world incident involving mismatched record counts and explain how the root cause was identified and resolved.
  • Proven experience integrating third-party REST APIs in production.
  • Experience handling pagination edge cases, rate and row limits, retries, and schema drift.
  • Experience with entity resolution or data matching involving messy, real-world text data.
  • Experience with a metrics or semantic layer such as Holistics AML/AQL, dbt Metrics, or LookML.
  • Working understanding of why non-additive measures cannot be reliably calculated from pre-aggregated rollups.
  • Strong SQL and Python skills, with the ability to own data pipelines end-to-end with minimal oversight.
  • Strong written and spoken English, with the ability to collaborate effectively with a US-based team asynchronously.
  • Healthcare data experience, including referrals, payer taxonomy, claims/eligibility, or other PHI-adjacent datasets.
  • Familiarity with HIPAA handling expectations.
  • Experience with voice-agent, call-center, or telephony/conversation data.
  • Familiarity with call transcripts and containment or outcome metrics.
  • Hands-on experience with Holistics, specifically AML/AQL modeling.
  • Experience with the broader Azure ecosystem beyond Event Hubs, including ADLS, ADF, and Key Vault.

Responsibilities

  • Build and maintain resilient ingestion pipelines for third-party vendor REST APIs.
  • Work primarily with voice-AI observability and telephony platforms.
  • Handle different pagination schemes, including offset/limit and page/cursor models.
  • Manage row and rate limits through time-windowing and adaptive bisection.
  • Implement robust schema-drift handling through contract and column-presence checks.
  • Ensure alerts are triggered when fields are renamed, moved, or removed rather than silently propagating null values.
  • Own streaming ingestion from Azure Event Hubs into Databricks using Structured Streaming and/or Auto Loader.
  • Manage checkpoints and offsets, watermarking, and at-least-once deduplication.
  • Perform source-parity reconciliation across ingested and production data.
  • Investigate row counts, dropped or duplicated events, late-arriving data, and schema mismatches.
  • Identify and resolve the root cause when ingested data does not match production sources.
  • Develop and maintain Delta Lake pipelines using a Medallion Architecture (Bronze Silver Gold).
  • Use Spark SQL and PySpark to build and maintain production data pipelines.
  • Implement idempotent MERGE upserts.
  • Work with Delta Live Tables and materialized-view constraints, including CREATE OR REFRESH and LIVE references.
  • Understand and manage differences between DLT and job execution contexts.
  • Build entity-resolution pipelines for dirty, free-text data.
  • Normalize practice, provider, and payer names using regex, canonical dictionaries, fuzzy matching, confidence-scored crosswalks, and override tables.
  • Maintain the semantic and metrics layer with rigorous metric definitions.
  • Define and maintain accurate denominators, data grain, and cohort boundaries.
  • Ensure the correct handling of non-additive aggregates, including medians and percentiles that cannot be reliably supported through aggregate-aware pre-aggregation.
  • Ensure every metric remains accurate and reproducible.
  • Instrument data quality across the entire pipeline.
  • Monitor data freshness, source parity, data contracts, and other critical quality checks.
  • Build alerting mechanisms that identify data issues before they reach dashboards.

Skills

  • Databricks
  • Spark
  • SQL
  • PySpark
  • Delta Lake
  • Delta Live Tables
  • Azure Event Hubs
  • Kafka
  • AWS Kinesis
  • dbt
  • Python
  • HIPAA
  • Azure
  • Azure Key Vault
  • Cursor

More jobs at HR POD - Hiring Talent Globally

All 12

Similar roles