NVIDIA AI Infrastructure & Kubernetes Platform Engineer (DGX Systems)
One IT
- Location
- US
- Workplace
- Remote
- Employment
- Contract
- Salary
- USD 90–100/hr
Posted 28d ago
The employer’s full description could not be read from their board. This is a summary of the posting — follow the apply link for the original.
Responsibilities
- Deploy and manage NVIDIA DGX BasePODs and SuperPODs for high-performance AI workloads
- Oversee DGX system lifecycle operations including provisioning, monitoring, firmware upgrades, and capacity planning
- Operate Base Command Manager to manage GPU clusters, schedule workloads, and integrate with MLOps tools
- Architect secure, scalable Kubernetes clusters optimized for GPU-accelerated workloads using the NVIDIA GPU Operator
- Apply CKA/CKAD/CKS expertise to develop, deploy, and secure AI applications on Kubernetes
- Implement CI/CD pipelines and GitOps methodologies for deploying and managing ML workflows
- Administer InfiniBand networks and BlueField DPUs using Unified Fabric Manager (UFM)
- Enable NVLink/NVSwitch performance across GPU nodes and tune fabric configurations for minimal latency and maximum throughput
- Use BlueField to offload storage, firewalling, and telemetry, strengthening AI workload security and performance
- Apply CKS best practices to secure containerized AI environments
- Configure runtime security, secrets management, network segmentation, and auditing across DPU-enhanced Kubernetes deployments
- Support zero-trust initiatives by enforcing workload identity, RBAC policies, and supply-chain integrity across AI container images and model artifacts
- Monitor GPU, CPU, and I/O performance using NVIDIA DCGM, Prometheus, Grafana, and Base Command APIs
- Tune system performance and model-training pipelines for cost-efficiency and throughput
- Build and maintain operational runbooks, incident-response playbooks, and SLA dashboards covering GPU utilization, thermal thresholds, and fabric health
Requirements
- Certified Kubernetes Administrator (CKA)
- Certified Kubernetes Application Developer (CKAD)
- Certified Kubernetes Security Specialist (CKS)
- NVIDIA Certified Associate: AI Infrastructure & Operations (NCA-AIIO)
- NVIDIA Certified Professional: AI Infrastructure (NCP-AII)
- NVIDIA Certified Professional: AI Operations (NCP-AIO)
- NVIDIA Certified Professional: AI Networking (NCP-AIN)
- DGX System, BasePOD, and SuperPOD administration
- BlueField DPU configuration and operations
- InfiniBand fabric and UFM management
- Base Command Manager for workload orchestration
- Kubernetes, Helm, and the NVIDIA GPU Operator
- DevOps tooling: Ansible, Terraform, GitOps, CI/CD pipelines
- Programming/scripting: Python, YAML, Bash
- Bachelor's degree in Computer Science, Engineering, or related field or equivalent hands-on experience
Preferred
- Kubeflow and broader MLOps pipeline experience
- Parallel/HPC storage: NFS, BeeGFS, Lustre
- Advanced networking: RoCE, RDMA, gRPC, and DPU offload tuning
Skills
- NVIDIA DGX
- Kubernetes
- NVIDIA GPU Operator
- InfiniBand
- BlueField DPU
- Unified Fabric Manager
- Base Command Manager
- NVLink
- NVSwitch
- Helm
- Ansible
- Terraform
- GitOps
- CI/CD
- Python
- YAML
- Bash
- NVIDIA DCGM
- Prometheus
- Grafana
- Kubeflow
- NFS
- BeeGFS
- Lustre
- RoCE
- RDMA
- gRPC
Similar roles
IT Support Specialist-III
Amperesand · Reno, Nevada, United States · today
Staff AI Engineering - Enterprise Architecture
American Express · Phoenix, AZ, United States · USD 144,250–256,250/yr · today
Staff AI Engineering - Enterprise Architecture
American Express · Phoenix, AZ, United States · USD 144,250–256,250/yr · today
Staff AI Infrastructure Engineer
Biohub · Redwood City, CA (Hybrid) · USD 241,000–331,000/yr · today
Senior Systems Engineer (Cloud Engineer)
UFCU Main · Austin, TX, United States · today
Field Service Technician II
Toshiba America Business Solutions · East Syracuse, NY, United States · today