JobHabor

Senior AI Systems Performance Engineer

SambaNova Systems

Location
US
Workplace
Onsite
Employment
Full Time
Salary
Apply on the employer’s site

Posted 3mo ago

The employer’s full description could not be read from their board. This is a summary of the posting — follow the apply link for the original.

Responsibilities

  • Optimize and scale foundation models on SambaNova's platform
  • Profile and enhance model performance across compiler, runtime, and hardware
  • Integrate advances in model architecture, quantization, scheduling, and memory optimization
  • Develop scalable and efficient end-to-end inference solutions
  • Identify performance bottlenecks and propose optimizations

Requirements

  • Bachelor's or higher degree in computer science, electrical engineering, or related field
  • 3+ years of experience in deep learning model development and performance optimization
  • 3+ years of experience in compiler, runtime, or kernel-level optimization
  • 3+ years of experience in software–hardware co-design or systems performance tuning
  • Proficiency in Python or C++
  • Strong foundations in algorithms, data structures, and numerical computing
  • Experience with PyTorch, TensorFlow, or JAX
  • Demonstrated ability to analyze and optimize performance in real-world ML pipelines

Preferred

  • Hands-on experience with LLM or multimodal model training and inference
  • Background in large-scale distributed training, continuous batching, and high-throughput inference systems
  • Familiarity with quantization, graph optimization, kernel fusion, and model partitioning
  • Experience with DeepSpeed, Megatron, vLLM, or TensorRT
  • Strong GPU programming skills (CUDA, Triton, or OpenCL)
  • Experience with cuDNN, cuBLAS, or similar libraries
  • Knowledge of memory hierarchy optimization, caching, and scheduling for large-scale model execution
  • Publication record or open-source contributions in ML systems or performance optimization

Skills

  • Python
  • C++
  • PyTorch
  • TensorFlow
  • JAX
  • CUDA
  • Triton
  • OpenCL
  • cuDNN
  • cuBLAS
  • DeepSpeed
  • Megatron
  • vLLM
  • TensorRT

Similar roles