JobHabor

Senior NVIDIA AI Factory / GPU Infrastructure Engineer

LiveMindz

Location
Remote
Workplace
Remote
Employment
Contract
Salary
Apply on the employer’s site

Posted 28d ago

Senior NVIDIA AI Factory / GPU Infrastructure Engineer NVIDIA H200

Position Overview

We are looking for a Senior NVIDIA AI Factory / GPU Infrastructure Engineer to design, deploy, configure, and productionize a large-scale NVIDIA H200 GPU environment for an enterprise AI Factory initiative.

The engineer will work on an environment consisting of 32 NVIDIA H200 GPUs across HPE Gen12 GPU servers, high-speed InfiniBand networking, NVIDIA AI Enterprise, GPU workload orchestration, multi-tenant inference, NVIDIA NIM, and token/API-based AI services.

This is a hands-on role requiring strong experience taking GPU infrastructure from hardware readiness through production AI workloads. Physical rack/stack and power activities will be handled by the on-site infrastructure partner.

Key Responsibilities

GPU Infrastructure & Cluster Deployment

  • Validate readiness of HPE DL380a Gen12 GPU servers with NVIDIA H200 GPUs.
  • Perform GPU, NVLink, PCIe, firmware, BIOS, BMC and hardware health validation.
  • Configure and maintain NVIDIA GPU drivers, CUDA, container runtime and supporting libraries.
  • Validate GPU topology and multi-GPU communication.
  • Configure and troubleshoot high-speed InfiniBand/NDR networking.
  • Configure and validate GPUDirect RDMA and GPU-to-GPU communication.
  • Perform NCCL and GPU communication/performance testing.
  • Integrate high-performance storage for AI models, datasets and inference workloads.
  • Establish validated firmware, OS, driver and CUDA baselines.

SLURM, MIG & GPU Resource Management

  • Deploy and configure SLURM for GPU workload scheduling and orchestration.
  • Configure NVIDIA Multi-Instance GPU (MIG) profiles where appropriate.
  • Implement GPU resource allocation, scheduling, quotas and workload isolation.
  • Support multi-user and multi-tenant GPU environments.
  • Monitor GPU utilization, memory, workload performance and cluster health.
  • Troubleshoot GPU scheduling, CUDA, NCCL, networking and performance issues.

AI Factory Architecture

Build two isolated, highly available GPU clusters

one internal and one public-facing.

  • Design production-grade GPU cluster architecture with security and workload isolation.
  • Implement high availability, monitoring, logging and operational controls.
  • Support capacity planning and GPU sizing for AI/LLM workloads.
  • Develop deployment runbooks, architecture documentation and operational procedures.

NVIDIA AI Enterprise & NIM

  • Deploy and configure NVIDIA AI Enterprise (NVAIE) components.
  • Deploy LLMs and AI models using NVIDIA NIM.
  • Configure scalable model-serving infrastructure.
  • Optimize inference workloads for NVIDIA H200 GPUs.
  • Perform model-serving benchmarking, capacity testing and performance tuning.
  • Integrate inference endpoints with enterprise API gateways.

Multi-Tenant Inference Platform

  • Design and implement a secure Inference-as-a-Service platform.
  • Configure API-based access to hosted AI/LLM models.
  • Implement tenant, application and model-level isolation.
  • Support API key management and secure service access.
  • Implement quotas, rate limits and GPU resource controls.
  • Enable token-based usage metering and consumption tracking.
  • Support model registry, prompt logging and AI governance controls.

Token Metering & API Gateway

  • Integrate inference services with an enterprise API gateway.
  • Capture input/output token consumption for LLM requests.
  • Implement token metering, quotas and rate limiting.
  • Enable per-tenant usage reporting and chargeback.
  • Integrate usage data with billing/payment platforms.
  • Support prepaid and postpaid AI consumption models.

Production Readiness

  • Conduct end-to-end infrastructure and inference testing.
  • Validate GPU performance, networking, storage and workload scheduling.
  • Perform HA/failover and resiliency testing.
  • Support security hardening and production-readiness assessments.
  • Establish monitoring, alerting and operational dashboards.
  • Define production acceptance criteria and SLA measurements.
  • Provide technical documentation and knowledge transfer to client teams.

Required Skills

  • 8+ years of Linux, infrastructure, HPC, cloud platform or systems engineering experience.
  • Strong hands-on experience with NVIDIA GPU infrastructure.
  • Experience deploying NVIDIA H100/H200 or comparable enterprise GPU platforms.

Strong knowledge of

  • NVIDIA CUDA
  • NVIDIA GPU Drivers
  • NVIDIA Container Toolkit
  • NCCL
  • GPUDirect RDMA
  • NVIDIA AI Enterprise
  • Hands-on experience with SLURM.
  • Experience with NVIDIA MIG and GPU resource partitioning.
  • Strong Linux administration and troubleshooting skills.
  • Experience with InfiniBand/RDMA high-performance networking.
  • Understanding of GPU topology, PCIe, NUMA, NVLink and multi-GPU communication.
  • Experience deploying containerized AI/ML workloads.
  • Experience with Docker and Kubernetes/container orchestration technologies.
  • Strong understanding of HA, monitoring, logging and production infrastructure practices.

AI/LLM Platform Experience

Candidates should have hands-on experience with several of the following:

  • NVIDIA NIM
  • NVIDIA AI Enterprise
  • LLM inference infrastructure
  • vLLM, Triton Inference Server or similar serving technologies
  • API gateways
  • Multi-tenant AI platforms
  • Token metering and usage tracking
  • Model registries
  • AI governance and guardrails
  • PrometheGrafana or equivalent observability platforms
  • Python/Bash automation

Preferred Experience

  • HPE ProLiant DL380/DL360 or similar enterprise GPU server platforms.
  • NVIDIA H100/H200 deployments.
  • Large-scale enterprise AI Factory implementations.
  • Production LLM serving and inference optimization.
  • Kubernetes GPU Operator.
  • Infrastructure automation using Ansible/Terraform.
  • Enterprise API gateway platforms.
  • FinOps/chargeback platforms.
  • Payment or billing-system integrations.
  • Telecom or other large regulated enterprise environments.
  • Security, data-sovereignty and compliance requirements.

Ideal Candidate

The ideal candidate is not solely an AI/ML developer or a traditional Linux administrator.We are looking for someone who understands the complete

GPU-to-API stack

HPE GPU Hardware Linux/Firmware NVIDIA Drivers/CUDA InfiniBand/RDMA SLURM/MIG GPU Clusters NVIDIA NIM LLM Inference API Gateway Multi-Tenancy Token Metering Production Operations

The candidate should be comfortable independently troubleshooting infrastructure from the GPU and network layer through the model-serving and API layer.

Engagement

Role

Senior NVIDIA AI Factory / GPU Infrastructure Engineer

Project

Enterprise AI Factory / GPU Platform

Initial Phase

Approximately 3 4 months

Overall Program

Potential long-term engagement

Environment

32 NVIDIA H200 GPUs across 4 GPU servers

Work Model

Primarily remote with coordination with an on-site infrastructure partner

Focus

GPU cluster deployment, SLURM/MIG, NVIDIA AI Enterprise, NIM, multi-tenant inference and production readiness

Skills

  • Linux
  • NVIDIA GPU
  • NVIDIA H100
  • NVIDIA H200
  • CUDA
  • NVIDIA GPU Drivers
  • NVIDIA Container Toolkit
  • NCCL
  • GPU Direct RDMA
  • NVIDIA AI Enterprise
  • SLURM
  • NVIDIA MIG
  • InfiniBand
  • RDMA
  • PCIe
  • NUMA
  • NVLink
  • Docker
  • Kubernetes
  • NVIDIA NIM
  • LLM inference
  • vLLM
  • Triton Inference Server
  • API gateways
  • Multi-tenant AI platforms
  • Token metering
  • Model registries
  • AI governance
  • Prometheus
  • Grafana
  • Python
  • Bash
  • HPE ProLiant DL380
  • HPE ProLiant DL360
  • Kubernetes GPU Operator
  • Ansible
  • Terraform
  • FinOps
  • chargeback platforms
  • payment systems
  • billing systems

Similar roles