JobHabor

Senior Site Reliability Engineer

Luma

Location
Location not stated
Workplace
Employment
Full Time
Salary
Apply on the employer’s site

Posted 1mo ago

Luma's mission is to build unified general intelligence that can generate, understand, and operate in the physical world, with a focus on multimodal intelligence and vision. The Senior Site Reliability Engineer will own and scale GPU infrastructure across on-premises and multi-cloud environments, ensuring reliable and performant training and inference clusters. The role involves Linux and kernel-level performance tuning, infrastructure automation, complex systems troubleshooting, re-architecture, and security compliance.

Artificial Intelligence (AI)

Media and Entertainment

Video

Foundational AI

Generative AI

Video Editing

check

H1B Sponsor Likelynote

check

Unicorn with $4B valuation

Insider Connection @Luma

2 email credits available today

note

Discover valuable connections within the company who might provide insights and potential referrals.

Get 3x more responses when you reach out via email instead of LinkedIn.

Beyond your network

K

K

I

Kate Tucker & 2 connections

From your previous company

F

F

F

& 2 connections

Previously@undefined and...

from your School

F

F

F

& 2 connections

@undefined and...

Find Any Email

Responsibilities

Take end-to-end ownership of production GPU clusters for training and inference across AWS and OCI, keeping them highly available and performant

Join critical re-architecture sessions to redesign systems for higher efficiency and scale

Tune Linux performance deeply, at the OS and kernel level

Build automation in Python, Go, or Bash to manage, monitor, and self-heal infrastructure without heavy toil

Serve as the final escalation for the hardest GPU, networking (InfiniBand/RDMA), and system failures, working with vendors like NVIDIA

Help achieve and maintain security certifications (SOC 2 Type 1 & 2, ISO) with strong infrastructure security practices

Qualification

check

Represents the skills you have

Find out how your skills align with this job's requirements. If anything seems off, you can easily click on the tags to select or unselect skills to reflect your actual expertise.

Linux

checkContainerized Systems

checkLow-Level Performance Debugging

checkTerraform

Apache Airflow

Ray

checkAWS

Oracle Cloud Infrastructure (OCI)

InfiniBand

RDMA

RoCE

SOC 2

ISO

NVIDIA DCGM

AMD ROCm

Required

5+ years as an SRE, production, or infrastructure engineer in a fast-paced, large-scale environment

Deep, hands-on Linux expertise, containerized systems, and low-level performance debugging

Working experience with Terraform, Airflow, and Ray

Strong experience with AWS or OCI

Practical experience with high-performance networking (InfiniBand, RDMA, or RoCE)

Working knowledge of security best practices and compliance frameworks like SOC 2 and ISO

Comfort in a less-structured, fast-paced environment

Preferred

Deep expertise with GPU tooling for NVIDIA and AMD (DCGM, ROCm)

Experience managing large-scale GPU clusters for AI/ML training or inference

Familiarity with Kubernetes or orchestration frameworks like Ray

Deep expertise in data pipelines and infrastructure

Skills

  • Linux
  • Python
  • Go
  • Bash
  • Terraform
  • Apache Airflow
  • Ray
  • AWS
  • OCI
  • InfiniBand
  • RDMA
  • RoCE
  • NVIDIA DCGM
  • AMD ROCm
  • Kubernetes
  • SOC 2
  • ISO

Similar roles