JobHabor

Senior Infrastructure Engineer

Earth Species Project

Location
Remote
Workplace
Remote
Employment
Full Time
Salary
USD 225,500–235,500/yr
Apply on the employer’s site

Posted 2mo ago

The employer’s full description could not be read from their board. This is a summary of the posting — follow the apply link for the original.

Responsibilities

  • Design and optimize high-performance data pipelines for distributed training and storage
  • Focus on low-level optimizations (latency, throughput, reliability, GPU usage)
  • Build monitoring and visualization tools for tracking data quality, pipeline performance, and experiments
  • Optimize distributed AI workloads for reliability, latency, and efficiency
  • Scope and supervise projects for interns, PhD students, and post-docs
  • Support recruiting efforts and help shape the infrastructure team

Requirements

  • 5+ years of backend or infrastructure engineering experience
  • Strong Python programming skills
  • Experience with distributed systems and cloud platforms (AWS, GCP, Azure)
  • Hands-on experience with containerization (Docker, Kubernetes)
  • Hands-on experience with infrastructure as code (Terraform)
  • Experience building or supporting ML/AI infrastructure in production
  • Experience with high-performance data tools (DuckDB, Apache Spark, Delta Lake)
  • GPU orchestration and large-scale model training experience
  • Familiarity with ML platforms (SageMaker, Vertex AI)
  • Familiarity with ML frameworks (PyTorch, JAX)
  • Experience mentoring junior engineers, interns, or researchers
  • Experience breaking down complex projects into manageable tasks
  • Experience participating in technical hiring processes and evaluating candidates

Preferred

  • Deep knowledge of training architectures, CUDA programming, or TPU optimization
  • Full-stack development experience with frameworks like React for building web applications
  • Experience managing HPC infrastructure with tools like Slurm or Kubernetes clusters
  • Background in monitoring stacks (Prometheus, Grafana) for ML pipeline observability

Skills

  • Python
  • AWS
  • GCP
  • Azure
  • Docker
  • Kubernetes
  • Terraform
  • DuckDB
  • Apache Spark
  • Delta Lake
  • SageMaker
  • Vertex AI
  • PyTorch
  • JAX
  • React
  • Slurm
  • Prometheus
  • Grafana
  • Arrow
  • LanceDB
  • BigQuery
  • CUDA

Similar roles