JobHabor

Staff Data Engineer

Arine
Location
San Francisco, US
Workplace
Hybrid
Employment
Full Time
Salary
USD 170,000–185,000/yr
Apply on the employer’s site

Posted 4mo ago

The employer’s full description could not be read from their board. This is a summary of the posting — follow the apply link for the original.

Responsibilities

  • Lead system design reviews
  • Conduct peer reviews
  • Architect and implement scalable data ingestion pipelines
  • Develop reusable, configuration-driven, containerized pipeline components
  • Collaborate with cross-functional teams on data requirements
  • Design and maintain data transformation pipelines using dbt
  • Build monitoring and alerting systems for data ingestion
  • Apply software engineering best practices to data infrastructure
  • Refactor existing ingestion processes
  • Provide technical guidance and mentorship
  • Promote best practices and coding standards
  • Champion AI-assisted development
  • Establish norms, workflows, and expectations for AI coding tools
  • Model the 'builder to reviewer' shift
  • Identify opportunities to automate repetitive engineering work using LLMs
  • Author and support technical documentation

Requirements

  • 10+ years in data engineering
  • Focus on large-scale data ingestion and infrastructure
  • Track record of building automated, production-grade ETL processes
  • Expert-level proficiency in Python
  • Expert-level proficiency in AWS services
  • Proficiency in DBT
  • Proficiency in SQL
  • Strong understanding of ETL/ELT frameworks
  • Strong understanding of distributed data processing
  • Hands-on experience building software with AI coding tools
  • Experience directing AI agents to generate complete solutions
  • Experience applying disciplined review and ownership of AI-generated code
  • Judgment to validate, test, and take accountability for AI-generated code
  • Experience or strong interest in integrating LLMs into engineering workflows
  • Proven ability to handle and process varied file types and formats
  • Experience with healthcare standards such as HL7, 834, 837, and NCPDP
  • Demonstrated success integrating and consolidating data from diverse source systems
  • Experience with EHR and claims systems integration
  • Experience with file-based and API integrations
  • Comfort working with large-scale datasets (10GB+)
  • Strong capability implementing incremental processing
  • Strong capability implementing change data capture (CDC) methodologies
  • Extensive background designing scalable data architectures in AWS environments
  • Solid grounding in software engineering principles
  • Experience with test-driven development
  • Experience with loose coupling
  • Experience with single responsibility
  • Experience with modular design
  • Hands-on familiarity with containerization (Docker, Kubernetes)
  • Proven ability to build configuration-driven systems
  • Passion for building new data infrastructure
  • Passion for continuously improving existing systems
  • Familiarity with healthcare data and regulatory environments (HIPAA)
  • Strong written and verbal communication skills
  • Comfort partnering across technical and non-technical stakeholders
  • Ability to pass a background check
  • Must live in and be eligible to work in the United States

Preferred

  • Experience with AI coding tools beyond autocomplete

Skills

  • Python
  • AWS
  • dbt
  • SQL
  • ETL
  • ELT
  • HL7
  • 834
  • 837
  • NCPDP
  • EHR
  • API
  • Docker
  • Kubernetes
  • LLMs
  • Claude Code
  • Cursor
  • Copilot
  • HIPAA

Similar roles