JobHabor

Senior Data Engineer

Plume

Location
US
Workplace
Remote
Employment
Full Time
Salary
USD 158,000–168,000/yr
Apply on the employer’s site

Posted 2mo ago

The employer’s full description could not be read from their board. This is a summary of the posting — follow the apply link for the original.

Responsibilities

  • Build and maintain production-grade data pipelines in cloud data warehouses
  • Design and develop dbt models across layers
  • Create and optimize Airflow DAGs for data workflow orchestration
  • Implement dimensional data models and data mart structures
  • Craft visualizations and dashboards in BI tools
  • Integrate healthcare data from various sources
  • Apply HIPAA-compliant data handling practices
  • Architect and implement RAG pipelines
  • Support MLOps workflows
  • Code review PRs from teammates
  • Collaborate with product managers
  • Monitor and triage pipeline and data quality failures
  • Document pipeline designs, data models, and technical decisions
  • Evaluate new tools and frameworks

Requirements

  • 5+ years of hands-on experience in data engineering, analytics engineering, or a closely related role
  • 2+ years of experience working within the healthcare industry
  • Working knowledge of HIPAA
  • Proven production experience with BigQuery, Snowflake, or Redshift
  • Strong hands-on experience with dbt
  • Deep experience with Apache Airflow
  • Demonstrated knowledge of dimensional data modeling
  • Hands-on experience delivering dashboards and reports in Looker, Power BI, Tableau, Qlik, etc.
  • Proficiency in Python for data pipeline development, API integrations, and automation
  • Practical exposure to RAG pipeline development and LLM integration using LangChain, LangGraph, or LlamaIndex
  • Hands-on exposure to MLOps concepts
  • Knowledge of CI/CD tooling for data and AI workloads
  • Strong understanding of data quality and governance principles
  • Excellent written and verbal communication skills
  • Ability to work independently

Preferred

  • Experience with real-time or streaming data pipelines using Kafka, Kinesis, or Pub/Sub
  • Knowledge of vector databases such as Pinecone, Weaviate, FAISS, or Chroma
  • Familiarity with responsible AI principles
  • Experience with data observability tools such as Monte Carlo, Bigeye, or Soda
  • Familiarity with data lakehouse patterns
  • Experience working toward or maintaining SOC2 or HITRUST certification
  • Familiarity with semantic layer tools
  • Experience with population health, revenue cycle, or clinical quality reporting datasets
  • Exposure to Kubernetes or containerized ML workloads

Skills

  • BigQuery
  • Snowflake
  • Redshift
  • dbt
  • Airflow
  • Looker
  • Power BI
  • Tableau
  • Qlik
  • Python
  • Pandas
  • PySpark
  • LangChain
  • LangGraph
  • LlamaIndex
  • GitHub Actions
  • Kafka
  • Kinesis
  • Pub/Sub
  • Pinecone
  • Weaviate
  • FAISS
  • Chroma
  • Delta Lake
  • Iceberg
  • Apache Hudi
  • LookML

Similar roles