Senior Site Reliability Engineer
Luma
- Location
- Location not stated
- Workplace
- —
- Employment
- Full Time
- Salary
- —
Posted 1mo ago
Luma's mission is to build unified general intelligence that can generate, understand, and operate in the physical world, with a focus on multimodal intelligence and vision. The Senior Site Reliability Engineer will own and scale GPU infrastructure across on-premises and multi-cloud environments, ensuring reliable and performant training and inference clusters. The role involves Linux and kernel-level performance tuning, infrastructure automation, complex systems troubleshooting, re-architecture, and security compliance.
Artificial Intelligence (AI)
Media and Entertainment
Video
Foundational AI
Generative AI
Video Editing
check
H1B Sponsor Likelynote
check
Unicorn with $4B valuation
Insider Connection @Luma
2 email credits available today
note
Discover valuable connections within the company who might provide insights and potential referrals.
Get 3x more responses when you reach out via email instead of LinkedIn.
Beyond your network
K
K
I
Kate Tucker & 2 connections
From your previous company
F
F
F
& 2 connections
Previously@undefined and...
from your School
F
F
F
& 2 connections
@undefined and...
Find Any Email
Responsibilities
Take end-to-end ownership of production GPU clusters for training and inference across AWS and OCI, keeping them highly available and performant
Join critical re-architecture sessions to redesign systems for higher efficiency and scale
Tune Linux performance deeply, at the OS and kernel level
Build automation in Python, Go, or Bash to manage, monitor, and self-heal infrastructure without heavy toil
Serve as the final escalation for the hardest GPU, networking (InfiniBand/RDMA), and system failures, working with vendors like NVIDIA
Help achieve and maintain security certifications (SOC 2 Type 1 & 2, ISO) with strong infrastructure security practices
Qualification
check
Represents the skills you have
Find out how your skills align with this job's requirements. If anything seems off, you can easily click on the tags to select or unselect skills to reflect your actual expertise.
Linux
checkContainerized Systems
checkLow-Level Performance Debugging
checkTerraform
Apache Airflow
Ray
checkAWS
Oracle Cloud Infrastructure (OCI)
InfiniBand
RDMA
RoCE
SOC 2
ISO
NVIDIA DCGM
AMD ROCm
Required
5+ years as an SRE, production, or infrastructure engineer in a fast-paced, large-scale environment
Deep, hands-on Linux expertise, containerized systems, and low-level performance debugging
Working experience with Terraform, Airflow, and Ray
Strong experience with AWS or OCI
Practical experience with high-performance networking (InfiniBand, RDMA, or RoCE)
Working knowledge of security best practices and compliance frameworks like SOC 2 and ISO
Comfort in a less-structured, fast-paced environment
Preferred
Deep expertise with GPU tooling for NVIDIA and AMD (DCGM, ROCm)
Experience managing large-scale GPU clusters for AI/ML training or inference
Familiarity with Kubernetes or orchestration frameworks like Ray
Deep expertise in data pipelines and infrastructure
Skills
- Linux
- Python
- Go
- Bash
- Terraform
- Apache Airflow
- Ray
- AWS
- OCI
- InfiniBand
- RDMA
- RoCE
- NVIDIA DCGM
- AMD ROCm
- Kubernetes
- SOC 2
- ISO
Similar roles
Sr Advanced Tech Product Owner
Honeywell · Bengaluru, Karnataka, India · today
Senior Business Engineer - Ads
Reddit · Remote · USD 180,200–252,300/yr · today
Machine Learning Engineer
Reddit · US · USD 185,800–303,400/yr · today
Senior Machine Learning Systems Engineer
Reddit · Global · USD 216,700–303,400/yr · today
Firmware Engineer
Anduril Industries · Costa Mesa, California, United States · USD 166,000–220,000/yr · today
Embedded Firmware Engineer
Anduril Industries · Costa Mesa, California, United States · USD 166,000–220,000/yr · today