JobHabor

Generative AI Inference Engineer

Stability AI

Location
United States
Workplace
Remote
Employment
Full Time
Salary
Apply on the employer’s site

Posted 2mo ago

The employer’s full description could not be read from their board. This is a summary of the posting — follow the apply link for the original.

Responsibilities

  • Lead design and development of customer-facing multi-modal ML inference systems
  • Build inference systems for next-generation models (optimization, tuning, deployment)
  • Partner with cloud providers for hosted Stability AI inference solutions
  • Be a strategic thought partner on driving business impact through ML
  • Bring new Stability models and pipelines into existence
  • Prototype and productionize inference platform improvements and new features

Requirements

  • 7+ years of productionizing machine learning systems, including inference pipeline development
  • Expert level knowledge on writing and running python services at scale
  • 5+ years working on python scientific stack, pyTorch
  • 5+ years working on at least one high-performance inference framework (e.g. Triton and TensorRT)
  • Deep understanding of Diffusion Architecture
  • Experience profiling and optimizing deep neural networks on Nvidia GPUs, using profiling tools such as NVIDIA Nsight
  • Experience with python-based image manipulation/encoding/decoding frameworks, such as OpenCV
  • Experience deploying to cloud orchestration systems such as Kubernetes
  • Experience deploying to cloud providers such as AWS, GCP, and Azure
  • Experience with Docker
  • Ability to rapidly prototype solutions and iterate on them with tight product deadlines
  • Experience with the open-source ML ecosystem (HuggingFace, W&B, etc.)

Preferred

  • Familiarity with workflow tools like ComfyUI

Skills

  • Python
  • PyTorch
  • Triton
  • TensorRT
  • Nvidia GPUs
  • NVIDIA Nsight
  • OpenCV
  • Kubernetes
  • AWS
  • GCP
  • Azure
  • Docker
  • HuggingFace
  • W&B

Similar roles