Staff Data Engineer
Arine- Location
- San Francisco, US
- Workplace
- Hybrid
- Employment
- Full Time
- Salary
- USD 170,000–185,000/yr
Posted 4mo ago
The employer’s full description could not be read from their board. This is a summary of the posting — follow the apply link for the original.
Responsibilities
- Lead system design reviews
- Conduct peer reviews
- Architect and implement scalable data ingestion pipelines
- Develop reusable, configuration-driven, containerized pipeline components
- Collaborate with cross-functional teams on data requirements
- Design and maintain data transformation pipelines using dbt
- Build monitoring and alerting systems for data ingestion
- Apply software engineering best practices to data infrastructure
- Refactor existing ingestion processes
- Provide technical guidance and mentorship
- Promote best practices and coding standards
- Champion AI-assisted development
- Establish norms, workflows, and expectations for AI coding tools
- Model the 'builder to reviewer' shift
- Identify opportunities to automate repetitive engineering work using LLMs
- Author and support technical documentation
Requirements
- 10+ years in data engineering
- Focus on large-scale data ingestion and infrastructure
- Track record of building automated, production-grade ETL processes
- Expert-level proficiency in Python
- Expert-level proficiency in AWS services
- Proficiency in DBT
- Proficiency in SQL
- Strong understanding of ETL/ELT frameworks
- Strong understanding of distributed data processing
- Hands-on experience building software with AI coding tools
- Experience directing AI agents to generate complete solutions
- Experience applying disciplined review and ownership of AI-generated code
- Judgment to validate, test, and take accountability for AI-generated code
- Experience or strong interest in integrating LLMs into engineering workflows
- Proven ability to handle and process varied file types and formats
- Experience with healthcare standards such as HL7, 834, 837, and NCPDP
- Demonstrated success integrating and consolidating data from diverse source systems
- Experience with EHR and claims systems integration
- Experience with file-based and API integrations
- Comfort working with large-scale datasets (10GB+)
- Strong capability implementing incremental processing
- Strong capability implementing change data capture (CDC) methodologies
- Extensive background designing scalable data architectures in AWS environments
- Solid grounding in software engineering principles
- Experience with test-driven development
- Experience with loose coupling
- Experience with single responsibility
- Experience with modular design
- Hands-on familiarity with containerization (Docker, Kubernetes)
- Proven ability to build configuration-driven systems
- Passion for building new data infrastructure
- Passion for continuously improving existing systems
- Familiarity with healthcare data and regulatory environments (HIPAA)
- Strong written and verbal communication skills
- Comfort partnering across technical and non-technical stakeholders
- Ability to pass a background check
- Must live in and be eligible to work in the United States
Preferred
- Experience with AI coding tools beyond autocomplete
Skills
- Python
- AWS
- dbt
- SQL
- ETL
- ELT
- HL7
- 834
- 837
- NCPDP
- EHR
- API
- Docker
- Kubernetes
- LLMs
- Claude Code
- Cursor
- Copilot
- HIPAA
Similar roles
Risk and Integrity Data Engineer
Evolution · Tbilisi, Tbilisi, Georgia · today
Staff Data Scientist, Imaging
Biohub · Redwood City, CA (Hybrid) · USD 214,000–294,800/yr · today
Senior/Staff Data Scientist
Render · SF · US/Canada · today
Senior Data Analyst, Avail
Realtor.com Careers · Austin, Texas, United States · today
Senior Analytics Engineer
Redwood Materials · San Francisco, California, United States · today
Data Engineering Manager
Redwood Materials · McCarran, NV · today