JobHabor

Senior Software Engineer I - AI Inference Data Plane

DigitalOcean

Location
US
Workplace
Remote
Employment
Full Time
Salary
USD 139,200–174,000/yr
Apply on the employer’s site

Posted 1mo ago

The employer’s full description could not be read from their board. This is a summary of the posting — follow the apply link for the original.

Responsibilities

  • Drive end-to-end design, development, and delivery of critical data plane components hosting large generative AI models
  • Architect and refine system design proposals for high-scale, multi-tenant AI inference cloud ecosystem
  • Implement and optimize distributed inference hosting using techniques like tensor/data parallelism, KV cache optimizations, and smart routing
  • Work cross-functionally with Product Managers, customer-facing teams, and other engineering teams
  • Build on Kubernetes-native distributed inference frameworks like llm-d (or alternatives such as NVIDIA Dynamo, Ray Serve, KServe)
  • Solve distributed-systems problems unique to LLM serving
  • Contribute upstream to llm-d, vLLM, and the inference gateway ecosystem
  • Coach and mentor junior engineers
  • Maintain and operate critical, high-scale services, utilizing observability tools and defining SLOs

Requirements

  • Hands-on experience hosting large language or multimodal models using inference engines like vLLM, SGLang, or TensorRT
  • Familiarity with distributed inference serving frameworks such as llm-d, NVIDIA Dynamo, or Ray Serve
  • Hands-on experience with vLLM or alternatives (SGLang, TensorRT-LLM, TGI, Modular MAX), including internals like continuous batching, paged attention, and prefix caching
  • Understanding of why cluster-scale serving is hard: KV-cache locality is partitioned across workers, naive round-robin routing destroys cache hit rates and tail latency, and disaggregated prefill/decode requires fast cross-pod KV transfer (e.g., NIXL)
  • Expert-level proficiency in GoLang or Python
  • Familiarity with gRPC
  • Proven experience shipping customer-facing software products and running critical services in a high-scale environment
  • Experience integrating and building with open-source software

Preferred

  • Merged contributions to vLLM, llm-d, SGLang, or similar projects

Skills

  • GoLang
  • Python
  • gRPC
  • Kubernetes
  • llm-d
  • NVIDIA Dynamo
  • Ray Serve
  • vLLM
  • SGLang
  • TensorRT
  • TensorRT-LLM
  • TGI
  • Modular MAX

Similar roles