ML Infrastructure Engineer, Training
Dyna Robotics- Location
- Redwood City, CA
- Workplace
- —
- Employment
- Full Time
- Salary
- —
Posted 5mo ago
Dyna Robotics builds general-purpose robots powered by a proprietary embodied AI foundation model with top-in-industry generalization and real-world performance. Already deployed with customers across multiple industries, our robots do commercial-grade work in the physical world. Our team comes from Google DeepMind, Meta, and Cruise, and we're backed by CRV, First Round, and other leading investors.
The Role
As a ML Training Infrastructure Engineer, you will architect and build the systems that turn our multi-cloud GPU fleet into a training engine our researchers love.
Your charter is singular and broad
own training infrastructure end-to-end so that every GPU is busy, every run is reproducible, and every researcher's next experiment is one command away.
What You’ll Do
- Scale Distributed Training: Architect and own the infrastructure for large-scale GPU clusters. You’ll implement sharding, activation checkpointing, and memory optimization (ZeRO, FSDP) to enable the training of massive multimodal models.
- Optimize Researcher Ergonomics: Build a research codebase and job scheduling system (Kubernetes/SLURM) that prioritizes fast iteration, automated retries, and seamless failure recovery.
- High-Performance Data Handling: Design high-throughput pipelines to ingest and transform terabytes of multimodal robot data (video, proprioception, 3D signals), ensuring dataloaders never starve the GPUs.
- Production Inference: Build low-latency inference pipelines for real-time robot control. You’ll apply quantization, distillation, and model compilation (TensorRT, Triton) to move models from the lab to the physical world.
- Deep Systems Profiling: Dive into the weeds of GPU utilization, I/O bottlenecks, and memory fragmentation to squeeze every bit of performance out of our expanding compute fleet.
What You’ll Bring
- 7+ Years of Engineering: With a track record of leading technical projects in high-performance computing (HPC) or ML infrastructure.
- ML Systems Mastery: Deep experience with PyTorch and distributed training frameworks (DeepSpeed, Accelerate). You understand the nuances of mixed precision and gradient accumulation.
- Infrastructure Expertise: Hands-on experience managing cloud GPU environments (GCP/AWS) and container orchestration (Kubernetes).
- Low-Level Intuition: A fundamental understanding of distributed systems, including race conditions, memory management, and NCCL/inter-node communication.
- Ownership Mindset: You don't just "deploy" code; you design, build, and operate systems end-to-end to unblock fast-moving research.
Bonus Points For
- Experience with Robotics Data Formats (MCAP, Protobuf) or multimodal models (VLAs).
- Deep ML systems experience: custom kernels (Triton), compilers, or runtime optimization.
- Experience as a founding or early-stage infrastructure hire.
At Dyna Robotics, we build technology for the real world, which requires a team as diverse as the environments our robots inhabit. We are an equal opportunity employer committed to technical rigor and mutual respect.
Don’t let a checklist stop you. Data shows that underrepresented groups often only apply if they meet 100% of the criteria. We value problem-solving and grit over keyword matching. If you’re passionate about robotics, no matter your discipline, we want to hear from you, even if you don't check every box.
Skills
- Machine Learning
- Kubernetes
- Slurm
- TensorRT
- Triton
- PyTorch
- GCP
- AWS
- Protobuf
More jobs at Dyna Robotics
All 8Exceptional Software Engineer
Dyna Robotics · Redwood City, CA · 1mo ago
Product Manager, Data Engine
Dyna Robotics · Redwood City, CA · 1mo ago
Annotation Operations Manager
Dyna Robotics · Redwood City, CA · 1mo ago
Applied Researcher - Deployment Intelligence & Continuous Learning
Dyna Robotics · Redwood City, CA · 1mo ago
Software Engineer, Applications
Dyna Robotics · Redwood City, CA · 5mo ago
Similar roles
IT Support Specialist-III
Amperesand · Reno, Nevada, United States · today
Staff AI Engineering - Enterprise Architecture
American Express · Phoenix, AZ, United States · USD 144,250–256,250/yr · today
Field Service Technician II
Toshiba America Business Solutions · East Syracuse, NY, United States · today
Field Service Technician I
Toshiba America Business Solutions · Rochester, NY, United States · today
Field Service Technician II
Toshiba America Business Solutions · Buffalo, NY, United States · today
Incentives Analyst
MicroStrategy · Tysons Corner, VIRGINIA, United States · USD 66,400–119,600/yr · today