Machine Learning Ops Engineer
Careforth- Location
- MN, United States
- Workplace
- Remote
- Employment
- Full Time
- Salary
- —
Posted 1mo ago
Position Summary
The ML Ops Engineer is a critical specialist within the Product & Technology Organization, responsible for the intersection of machine learning, software engineering, and platform operations. You will design and maintain the infrastructure required to scale ML models across clinical risk intelligence, caregiver risk scoring, composite risk trajectory, NLP signal capture, LLM-powered enablement tools, and insights reporting — from research through reliable production.
The ideal candidate has deep experience with distributed systems, containerization, model lifecycle governance, and automated ML pipelines, and thrives collaborating with Data Scientists and Data Engineers to ensure models are deployable, monitorable, HIPAA-compliant, and continuously improving.
What You Will Do
- Design and implement automated ML pipelines for model training, evaluation, and deployment using MLflow, Databricks Workflows, and AWS SageMaker Pipelines; own model registry governance including versioning, promotion, and retirement.
- Build and manage scalable model serving infrastructure (REST APIs, WebSocket APIs) using Docker and Kubernetes/EKS for real-time and batch scoring across all risk and enablement model domains.
- Architect and operate a sub-model orchestration layer, score computation service, score history, and audit logging compliant with HIPAA requirements.
- Design and maintain feature store architecture, temporal feature computation, and data versioning to ensure training-serving consistency across all ML domains.
- Implement and govern LLM API integrations (Claude/GPT via Bedrock) including prompt versioning, rate limiting, cost tracking, response logging, and output guardrails.
- Build production monitoring and alerting for data drift, model decay, scoring latency, and pipeline failures; implement A/B testing infrastructure for controlled model rollouts.
- Automate ML infrastructure provisioning using Terraform or AWS CloudFormation; manage secrets, access controls, PHI redaction, and HIPAA-compliant data handling across all services.
- Help define and lead the enterprise MLOps platform strategy, drive adoption of core tooling, and mentor junior engineers on operational standards and best practices.
- Establish CI/CD and Continuous Training (CT) workflows to enable rapid, safe, and auditable ML iteration across all project domains.
- Perform other duties and special projects as assigned.
What You Will Bring
Education
- Bachelor's or Master's Degree in Computer Science, Software Engineering, or a related technical field.
Experience
- 7+ years of professional experience in DevOps, Data Engineering, or ML Engineering, with at least 4 years focused on Machine Learning operations.
- Proven track record of owning and operating production ML systems including model serving, monitoring, and lifecycle governance.
- Experience in healthcare or regulated data environments; familiarity with HIPAA technical safeguards required.
- Experience mentoring engineers and contributing to platform standards and technical roadmaps.
Technical Skills
Expert-level containerization and orchestration
Docker, Kubernetes/EKS; Infrastructure-as-Code with Terraform or AWS CloudFormation.
Deep experience with ML lifecycle tooling
MLflow, Databricks (Delta Lake, Unity Catalog, Workflows, Spark, Vector Search), and AWS SageMaker.
Strong AWS proficiency
S3, Lambda, API Gateway, SageMaker, Bedrock, Redshift, Athena, DynamoDB, Glue, Step Functions, Kinesis, Transcribe, Comprehend Medical, Secrets Manager.
- Experience with stream processing (Kafka/Kinesis), event-driven pipeline design, and feature store architecture.
- Familiarity with LLM API integration, prompt versioning, RAG infrastructure, and LLM quality and cost governance.
- Strong Python proficiency; Bash scripting; Go or Java a plus.
Soft Skills
- Clear communicator able to translate operational constraints into actionable guidance for Data Scientists and product teams.
- Collaborative, detail-oriented, and committed to reproducibility, HIPAA audit-readiness, and operational excellence.
- Self-directed and intellectually curious; proactive in evaluating and adopting emerging MLOps tooling.
You'll Benefit From
At Careforth your well-being matters. With flexible schedules, a remote-first culture, and a nationally recognized wellness program, our benefits are designed to help you thrive, both professionally and personally.
Discover how we invest in you
https://careforth.com/careers/#benefits
The pay range for this position is $111,000 - $175,000. The actual wage offered may be lower or higher depending on budget and candidate experience, knowledge, skills, qualifications, and geographic location.
Skills
- Machine Learning
- NLP
- LLM
- HIPAA
- MLflow
- Databricks
- SageMaker
- WebSockets
- Docker
- Kubernetes
- EKS
- Temporal
- Anthropic Claude
- GPT
- Bedrock
- Terraform
- AWS CloudFormation
- MLOps
- Delta Lake
- Unity Catalog
- Spark
- AWS
- S3
- AWS Lambda
- API Gateway
- Redshift
- Athena
- DynamoDB
- AWS Glue
- AWS Step Functions
- AWS Kinesis
- Comprehend
- Secrets Manager
- Kafka
- Retrieval-Augmented Generation
- Python
- Bash
- Go
- Java
More jobs at Careforth
All 7Untitled role
Careforth · Location not stated · 25d ago
Senior Software Engineer – Full Stack
Careforth · US · USD 111,064–166,596/yr · 1mo ago
Data Scientist II
Careforth · MN, United States · 1mo ago
Manager, Software Engineering and Operations
Careforth · MA, United States · 1mo ago
Manager, Software Engineering
Careforth · CT, United States · 1mo ago
Similar roles
Senior Software Engineer
Redwood Materials · remote · Nevada · USD 180,000–237,500/yr · today
Embedded Software Engineer – Power Electronics, Energy Storage
Redwood Materials · San Francisco, California, United States · USD 180,000–237,500/yr · today
Software Engineer - ML/Computer Vision (Battery Sorting)
Redwood Materials · McCarran, NV · San Francisco, California, United States · USD 152,500–200,000/yr · today
Principal Software Engineer, Core Infrastructure
Oracle · Nashville, TN, United States · today
Senior Software Engineer
Advantest · Lake Forest, CA, United States · today
Senior Frontend Engineer, Ads Creative
Reddit · US · USD 190,800–267,100/yr · today