LLM Inference & GPU Systems Consultant
Delan Associates- Location
- Charlotte, North Carolina, United States
- Workplace
- —
- Employment
- Contract
- Salary
- —
Posted 1mo ago
Job Title: LLM Inference & GPU Systems Consultant
Location: Charlotte, NC (Onsite)
Duration: 6+ Months
Must be onsite at client in Charlotte, NC at least 3 days/week
Role Overview
We are seeking an AI Infrastructure Runtime Engineer to build and maintain large-scale on-prem LLM infrastructure. This is an enterprise private GenAI environment running on NVIDIA H200 GPU clusters and an OpenShift AI deployment ecosystem. You will manage production inference internally, including self-hosting open-source LLMs like Llama. We are focused exclusively on inferencing; this role involves no model training infrastructure or fine-tuning pipelines.
Key Responsibilities
NVIDIA GPU Runtime Optimization
Drive extreme runtime efficiency and optimization for the token generation pipeline. Specifically manage prefill/decode optimization and KV cache management.
Inference Serving
Deploy and manage inference engines including vLLM and TensorRT-LLM.
Hardware Utilization
Optimize GPU throughput tuning, batching strategies, and latency optimization. Manage workload orchestration using RunAI and Kubernetes GPU orchestration.
Model Lifecycle Management
Oversee the complete Hugging Face model lifecycle, including model onboarding, deployment, and retirement.
Platform Operations
Operate and maintain the OpenShift AI ecosystem as the primary container platform for GenAI workloads.
Required Qualifications
8+ years experience working as an LLM Systems Engineer or AI Infrastructure Runtime Engineer.
8+ years hands-on experience with NVIDIA H200 clusters and runtime optimization techniques (KV Cache, prefill/decode).
Proficiency in OpenShift AI and GPU orchestration tools like RunAI.
Strong experience with modern inference frameworks, specifically vLLM and TensorRT-LLM.
Proven track record managing the Hugging Face deployment lifecycle.
Skills
- LLM
- Generative AI
- OpenShift
- Llama
- vLLM
- TensorRT
- Kubernetes
- Hugging Face
More jobs at Delan Associates
All 199Curam Developer
Delan Associates · New York City, New York, United States · 2d ago
Scrum Master
Delan Associates · New York City, New York, United States · 2d ago
Mobile Developere
Delan Associates · New York City, New York, United States · 2d ago
Computer Specialist
Delan Associates · New York City, New York, United States · 2d ago
Calypso Engineer / DevOps
Delan Associates · Atlanta, GA/ Charlotte, NC/ NYC, NY, Georgia, United States · 2d ago
Similar roles
Lead Software Engineer - City
Okc · Oklahoma City, OK, United States · USD 40–60/hr · today
Senior Software Engineer
Redwood Materials · remote · Nevada · USD 180,000–237,500/yr · today
Embedded Software Engineer – Power Electronics, Energy Storage
Redwood Materials · San Francisco, California, United States · USD 180,000–237,500/yr · today
Senior Software Engineer (Reston or Cambridge)
Akamai · United States · USD 121,400–218,600/yr · today
Software Engineer
Realtor.com Careers · Austin · today
Software Engineer - ML/Computer Vision (Battery Sorting)
Redwood Materials · McCarran, NV · San Francisco, California, United States · USD 152,500–200,000/yr · today