Senior Infrastructure Engineer
Earth Species Project
- Location
- Remote
- Workplace
- Remote
- Employment
- Full Time
- Salary
- USD 225,500–235,500/yr
Posted 2mo ago
The employer’s full description could not be read from their board. This is a summary of the posting — follow the apply link for the original.
Responsibilities
- Design and optimize high-performance data pipelines for distributed training and storage
- Focus on low-level optimizations (latency, throughput, reliability, GPU usage)
- Build monitoring and visualization tools for tracking data quality, pipeline performance, and experiments
- Optimize distributed AI workloads for reliability, latency, and efficiency
- Scope and supervise projects for interns, PhD students, and post-docs
- Support recruiting efforts and help shape the infrastructure team
Requirements
- 5+ years of backend or infrastructure engineering experience
- Strong Python programming skills
- Experience with distributed systems and cloud platforms (AWS, GCP, Azure)
- Hands-on experience with containerization (Docker, Kubernetes)
- Hands-on experience with infrastructure as code (Terraform)
- Experience building or supporting ML/AI infrastructure in production
- Experience with high-performance data tools (DuckDB, Apache Spark, Delta Lake)
- GPU orchestration and large-scale model training experience
- Familiarity with ML platforms (SageMaker, Vertex AI)
- Familiarity with ML frameworks (PyTorch, JAX)
- Experience mentoring junior engineers, interns, or researchers
- Experience breaking down complex projects into manageable tasks
- Experience participating in technical hiring processes and evaluating candidates
Preferred
- Deep knowledge of training architectures, CUDA programming, or TPU optimization
- Full-stack development experience with frameworks like React for building web applications
- Experience managing HPC infrastructure with tools like Slurm or Kubernetes clusters
- Background in monitoring stacks (Prometheus, Grafana) for ML pipeline observability
Skills
- Python
- AWS
- GCP
- Azure
- Docker
- Kubernetes
- Terraform
- DuckDB
- Apache Spark
- Delta Lake
- SageMaker
- Vertex AI
- PyTorch
- JAX
- React
- Slurm
- Prometheus
- Grafana
- Arrow
- LanceDB
- BigQuery
- CUDA
Similar roles
SOFTWARE ENGINEERING MANAGER
NMDC Career site · Abu Dhabi, United Arab Emirates · United Arab Emirates · today
Forward Deployed Engineer, Hebrew speaker
Cloudflare · Hybrid · today
Software Engineer
Cloudflare · Location not stated · today
Software Engineer in Build Engineering
Graphcore · Bristol, UK · today
Software Engineer in Build Engineering
Graphcore · Cambridge, UK · today
Software Engineer in Build Engineering
Graphcore · London, UK · today