ML Platform Engineer
Bright Vision Technologies
- Location
- Houston
- Workplace
- Remote
- Employment
- Full Time
- Salary
- USD 100,000–150,000/yr
Posted 2mo ago
The employer’s full description could not be read from their board. This is a summary of the posting — follow the apply link for the original.
Responsibilities
- Design and operate model serving platforms
- Optimize inference performance
- Implement multi-tenant routing, rate limiting, and QoS policies
- Build autoscaling and capacity management systems
- Tune GPU utilization, memory management, and KV cache strategies
- Integrate model serving with API gateways, identity systems, and observability platforms
- Implement caching, prompt deduplication, and response reuse
- Drive end-to-end observability
- Develop deployment workflows
- Operate incident response for high-availability AI services
- Collaborate with ML and product teams
- Implement security controls
- Document operational procedures, performance characteristics, and tuning guidance
- Stay current with AI serving research
Requirements
- Bachelor’s or Master’s degree in Computer Science or related field
- Six or more years of experience in distributed systems, infrastructure, or ML platform engineering
- Strong proficiency in Python
- Strong proficiency in a systems language such as Go, Rust, or C++
- Deep experience operating high-throughput, low-latency services in production
- Hands-on experience with LLM or large model inference frameworks such as vLLM or TensorRT-LLM
- Strong understanding of GPU architecture, memory hierarchies, and accelerator utilization
- Familiarity with Kubernetes
- Familiarity with autoscaling
- Familiarity with modern cloud platforms
- Experience with observability stacks including metrics, tracing, and structured logging
- Solid grounding in performance engineering and capacity planning
- Strong communication skills
- Strong incident response skills
Preferred
- Open-source contributions to model serving infrastructure
- Experience with multi-region or globally distributed AI serving
- Familiarity with model quantization, distillation, and compression techniques
- Exposure to FinOps for AI workloads and cost-efficient serving design
- Experience supporting external-facing AI APIs at scale
Skills
- Python
- Go
- Rust
- C++
- vLLM
- TensorRT-LLM
- Kubernetes
Similar roles
Site Reliability Engineer (FedRAMP / Security)
Coralogix · New York, NY, United States · USD 170,000–350,000/yr · today
Senior Staff Product Manager, Conversational AI
ServiceNow · Santa Clara, California, United States · USD 190,900–334,100/yr · today
Senior Business Engineer - Ads
Reddit · Remote · USD 180,200–252,300/yr · today
Machine Learning Engineer
Reddit · US · USD 185,800–303,400/yr · today
Backend Software Engineer
ClearlyRated · Portland, OR, United States · USD 90,000–120,000/yr · today
Firmware Engineer
Anduril Industries · Costa Mesa, California, United States · USD 166,000–220,000/yr · today