JobHabor

ML Platform Engineer

Bright Vision Technologies

Location
Houston
Workplace
Remote
Employment
Full Time
Salary
USD 100,000–150,000/yr
Apply on the employer’s site

Posted 2mo ago

The employer’s full description could not be read from their board. This is a summary of the posting — follow the apply link for the original.

Responsibilities

  • Design and operate model serving platforms
  • Optimize inference performance
  • Implement multi-tenant routing, rate limiting, and QoS policies
  • Build autoscaling and capacity management systems
  • Tune GPU utilization, memory management, and KV cache strategies
  • Integrate model serving with API gateways, identity systems, and observability platforms
  • Implement caching, prompt deduplication, and response reuse
  • Drive end-to-end observability
  • Develop deployment workflows
  • Operate incident response for high-availability AI services
  • Collaborate with ML and product teams
  • Implement security controls
  • Document operational procedures, performance characteristics, and tuning guidance
  • Stay current with AI serving research

Requirements

  • Bachelor’s or Master’s degree in Computer Science or related field
  • Six or more years of experience in distributed systems, infrastructure, or ML platform engineering
  • Strong proficiency in Python
  • Strong proficiency in a systems language such as Go, Rust, or C++
  • Deep experience operating high-throughput, low-latency services in production
  • Hands-on experience with LLM or large model inference frameworks such as vLLM or TensorRT-LLM
  • Strong understanding of GPU architecture, memory hierarchies, and accelerator utilization
  • Familiarity with Kubernetes
  • Familiarity with autoscaling
  • Familiarity with modern cloud platforms
  • Experience with observability stacks including metrics, tracing, and structured logging
  • Solid grounding in performance engineering and capacity planning
  • Strong communication skills
  • Strong incident response skills

Preferred

  • Open-source contributions to model serving infrastructure
  • Experience with multi-region or globally distributed AI serving
  • Familiarity with model quantization, distillation, and compression techniques
  • Exposure to FinOps for AI workloads and cost-efficient serving design
  • Experience supporting external-facing AI APIs at scale

Skills

  • Python
  • Go
  • Rust
  • C++
  • vLLM
  • TensorRT-LLM
  • Kubernetes

Similar roles