JobHabor

Senior Data Engineer

Location
United States
Workplace
Remote
Employment
Full Time
Salary
Apply on the employer’s site

Posted 27d ago

The employer’s full description could not be read from their board. This is a summary of the posting — follow the apply link for the original.

Responsibilities

  • Design, build, and maintain scalable data pipelines and ETL/ELT processes
  • Develop data engineering solutions using Python, SQL, and Databricks
  • Build and optimize data pipelines within the Databricks Lakehouse environment
  • Work with large, complex datasets from multiple sources
  • Partner with engineering teams to extract, integrate, and modernize data flows from legacy systems, including IBMi/AS400 environments
  • Contribute to Databricks implementation project and support data platform readiness for large scale data conversion
  • Design and implement change data capture (CDC) patterns
  • Develop and maintain data models and curate datasets for analytics, reporting, and downstream applications
  • Implement data quality, validation, monitoring, and error-handling processes
  • Optimize data pipelines and queries for performance, scalability, reliability, and cost
  • Partner with data architects, analysts, data scientists, application teams, and business stakeholders
  • Collaborate with the AI governance group to prepare and structure data for AI/ML pilot programs
  • Participate in the design and evolution of MMG's modern data architecture and engineering practices
  • Establish and promote standards for code quality, testing, documentation, version control, and deployment
  • Troubleshoot complex data and pipeline issues and provide sustainable solutions
  • Contribute to CI/CD and automated deployment practices for data engineering workloads
  • Mentor other engineers and contribute to the growth of MMG's data engineering capabilities

Requirements

  • 5+ years of professional experience in data engineering, software engineering, or a related technical discipline
  • Strong professional experience with Python
  • Strong experience with Databricks, including developing and optimizing data pipelines and workloads
  • Strong SQL skills and experience working with relational and/or analytical databases
  • Experience designing and implementing ETL/ELT pipelines and data integration solutions
  • Experience working with cloud-based data platforms and modern data architectures (Azure preferred)
  • Strong understanding of data modeling, data warehousing, and data lake/lakehouse concepts including medallion / multi-hop architectures
  • Experience with Git and modern software development practices
  • Experience with automated testing, deployment, and CI/CD practices
  • Experience establishing version control, testing, review and deployment automation for data
  • Bachelor's degree in Computer Science, Information Systems, Engineering, or a related field, or equivalent professional experience

Preferred

  • Experience working in the insurance or financial services industry
  • Experience integrating data from legacy systems, particularly IBMi/AS400 environments
  • Experience with Azure AI Foundry or similar AI/ML platform tooling
  • Experience with Apache Spark / PySpark
  • Experience with Databricks SQL and Delta Lake
  • Experience with data orchestration tools such as Azure Data Factory, Airflow, or similar technologies
  • Experience with APIs and event-driven or real-time data integration
  • Experience implementing change data capture (CDC) patterns for enterprise data integration
  • Experience with insurance policy administration platform data or large-scale platform data conversions
  • Experience with data governance, metadata, lineage, and data quality frameworks
  • Experience working with enterprise data warehouses and dimensional modeling
  • Experience with Infrastructure as Code and cloud automation
  • Experience using AI-assisted development tools to improve engineering productivity

Skills

  • Python
  • Databricks
  • SQL
  • ETL
  • ELT
  • Data Modeling
  • Git
  • Azure
  • PySpark
  • Delta Lake
  • Azure Data Factory
  • Airflow
  • Azure AI Foundry
  • Databricks SQL
  • APIs
  • CDC
  • Data Governance
  • IBMi
  • AS400
  • Apache Spark

Similar roles