JobHabor

NVIDIA AI Infrastructure & Kubernetes Platform Engineer (DGX Systems)

One IT

Location
US
Workplace
Remote
Employment
Contract
Salary
USD 90–100/hr
Apply on the employer’s site

Posted 28d ago

The employer’s full description could not be read from their board. This is a summary of the posting — follow the apply link for the original.

Responsibilities

  • Deploy and manage NVIDIA DGX BasePODs and SuperPODs for high-performance AI workloads
  • Oversee DGX system lifecycle operations including provisioning, monitoring, firmware upgrades, and capacity planning
  • Operate Base Command Manager to manage GPU clusters, schedule workloads, and integrate with MLOps tools
  • Architect secure, scalable Kubernetes clusters optimized for GPU-accelerated workloads using the NVIDIA GPU Operator
  • Apply CKA/CKAD/CKS expertise to develop, deploy, and secure AI applications on Kubernetes
  • Implement CI/CD pipelines and GitOps methodologies for deploying and managing ML workflows
  • Administer InfiniBand networks and BlueField DPUs using Unified Fabric Manager (UFM)
  • Enable NVLink/NVSwitch performance across GPU nodes and tune fabric configurations for minimal latency and maximum throughput
  • Use BlueField to offload storage, firewalling, and telemetry, strengthening AI workload security and performance
  • Apply CKS best practices to secure containerized AI environments
  • Configure runtime security, secrets management, network segmentation, and auditing across DPU-enhanced Kubernetes deployments
  • Support zero-trust initiatives by enforcing workload identity, RBAC policies, and supply-chain integrity across AI container images and model artifacts
  • Monitor GPU, CPU, and I/O performance using NVIDIA DCGM, Prometheus, Grafana, and Base Command APIs
  • Tune system performance and model-training pipelines for cost-efficiency and throughput
  • Build and maintain operational runbooks, incident-response playbooks, and SLA dashboards covering GPU utilization, thermal thresholds, and fabric health

Requirements

  • Certified Kubernetes Administrator (CKA)
  • Certified Kubernetes Application Developer (CKAD)
  • Certified Kubernetes Security Specialist (CKS)
  • NVIDIA Certified Associate: AI Infrastructure & Operations (NCA-AIIO)
  • NVIDIA Certified Professional: AI Infrastructure (NCP-AII)
  • NVIDIA Certified Professional: AI Operations (NCP-AIO)
  • NVIDIA Certified Professional: AI Networking (NCP-AIN)
  • DGX System, BasePOD, and SuperPOD administration
  • BlueField DPU configuration and operations
  • InfiniBand fabric and UFM management
  • Base Command Manager for workload orchestration
  • Kubernetes, Helm, and the NVIDIA GPU Operator
  • DevOps tooling: Ansible, Terraform, GitOps, CI/CD pipelines
  • Programming/scripting: Python, YAML, Bash
  • Bachelor's degree in Computer Science, Engineering, or related field or equivalent hands-on experience

Preferred

  • Kubeflow and broader MLOps pipeline experience
  • Parallel/HPC storage: NFS, BeeGFS, Lustre
  • Advanced networking: RoCE, RDMA, gRPC, and DPU offload tuning

Skills

  • NVIDIA DGX
  • Kubernetes
  • NVIDIA GPU Operator
  • InfiniBand
  • BlueField DPU
  • Unified Fabric Manager
  • Base Command Manager
  • NVLink
  • NVSwitch
  • Helm
  • Ansible
  • Terraform
  • GitOps
  • CI/CD
  • Python
  • YAML
  • Bash
  • NVIDIA DCGM
  • Prometheus
  • Grafana
  • Kubeflow
  • NFS
  • BeeGFS
  • Lustre
  • RoCE
  • RDMA
  • gRPC

Similar roles