JobHabor

AI Systems Performance Specialist

Bright Vision Technologies

Location
United States
Workplace
Remote
Employment
Full Time
Salary
USD 130,000–180,000/yr
Apply on the employer’s site

Posted 28d ago

AI Systems Performance Specialist @ Bright Vision Technologies | Jobright.ai

AI Systems Performance Specialist jobs in United States

Overview

Company

This job has closed.

APPLY to similar jobs

Bright Vision Technologies · 2 weeks ago

AI Systems Performance Specialist

United States

Full-time

Remote

Lead/Staff

$130K/yr - $180K/yr

10+ years exp

Bright Vision Technologies is a technology consulting and software development company delivering cloud, AI, data, and enterprise solutions across the United States. The company is seeking an experienced AI Systems Performance Specialist to optimize AI training and inference workloads for performance, scalability, reliability, and cost efficiency, while leading optimization initiatives across enterprise-scale AI platforms.

Artificial Intelligence (AI)Cyber SecurityInformation TechnologySoftware

No H1B

Responsibilities

Optimize AI training and inference pipelines for maximum throughput, low latency, scalability, and infrastructure efficiency

Analyze and improve GPU utilization, memory management, kernel execution, and multi-GPU performance across production AI workloads

Design and implement optimization techniques including quantization, pruning, mixed precision, batching, caching, speculative decoding, and model parallelism

Profile AI applications using industry-standard performance analysis tools and identify bottlenecks across compute, memory, networking, and storage

Optimize distributed training and inference using NCCL, DeepSpeed, PyTorch Distributed, Ray, MPI, or similar distributed computing frameworks

Collaborate with AI researchers, ML engineers, platform engineers, and infrastructure teams to improve model performance and production reliability

Build automated benchmarking frameworks, performance dashboards, monitoring solutions, and regression testing pipelines

Evaluate emerging AI hardware, GPU architectures, inference frameworks, and optimization technologies to improve enterprise AI capabilities

Drive AI infrastructure cost optimization through efficient resource utilization, cloud optimization, and FinOps best practices

Mentor engineering teams and provide technical leadership on AI systems architecture, GPU optimization, and performance engineering

Qualification

PythonC++CUDAGPU OptimizationDistributed TrainingLarge Language Model InferenceDeep Learning FrameworksAI InfrastructureHigh-Performance ComputingPerformance EngineeringNVIDIA Nsight SystemsPyTorch ProfilerCloud Computing AWSCloud Computing Microsoft AzureCloud ComputingCloud Computing Google Cloud Platform

Required

**Experience Required

** **10+ Years**

**Sponsorship

** U.S. Citizens, Green Card Holders, EAD Holders, and H-1B transfer candidates are encouraged to apply. We are unable to sponsor new H-1B visa petitions for this position

  • Bachelor's or Master's degree in Computer Science, Computer Engineering, Electrical Engineering, Artificial Intelligence, or a related technical discipline
  • **10+ years of professional experience** in performance engineering, AI infrastructure, machine learning systems, High-Performance Computing (HPC), or distributed computing
  • Expert-level programming skills in **Python** and **C++**
  • Extensive experience optimizing GPU-accelerated AI workloads using **CUDA**, distributed training frameworks, and modern deep learning libraries
  • Strong knowledge of Large Language Models (LLMs), deep learning frameworks, model serving, and production AI inference
  • Hands-on experience with profiling tools such as **NVIDIA Nsight Systems, Nsight Compute, PyTorch Profiler, TensorBoard, or similar performance analysis tools**
  • Experience deploying and optimizing AI workloads on **AWS, Microsoft Azure, or Google Cloud Platform (GCP)**
  • Strong understanding of distributed systems, networking, storage optimization, and AI infrastructure architecture
  • Excellent analytical, troubleshooting, communication, and technical leadership skills

Preferred

  • Experience optimizing **production-scale LLM inference** and serving large foundation models
  • Hands-on experience with **vLLM**, **TensorRT-LLM**, **DeepSpeed**, **Triton Inference Server**, **CUTLASS**, **FasterTransformer**, or similar AI optimization frameworks
  • Knowledge of model compression, KV cache optimization, speculative decoding, and advanced inference optimization techniques
  • Experience implementing **FinOps** strategies for AI infrastructure cost optimization and resource management
  • Contributions to AI systems research, open-source AI infrastructure projects, patents, or technical publications
  • Familiarity with emerging AI accelerator technologies, including **AMD ROCm**, **Intel oneAPI**, or custom AI hardware

Benefits

100% Remote (Continental United States)

Full-time, Direct W2

Tremendous career growth potential

Company

Bright Vision Technologies

AI-powered talent intelligence and enterprise automation platform.

Founded in 2020

Bridgewater, New Jersey, USA

51-200 employees

https://bvteck.com

Funding

Current Stage

Growth Stage

Company data provided by crunchbase

Skills

  • Python
  • C++
  • CUDA
  • GPU Optimization
  • Distributed Training
  • Large Language Models
  • Deep Learning
  • AI Infrastructure
  • High-Performance Computing
  • Performance Engineering
  • NVIDIA Nsight Systems
  • PyTorch Profiler
  • AWS
  • Microsoft Azure
  • Google Cloud Platform
  • NCCL
  • DeepSpeed
  • PyTorch
  • Ray
  • MPI
  • vLLM
  • TensorRT-LLM
  • Triton Inference Server
  • CUTLASS
  • FasterTransformer
  • FinOps
  • AMD ROCm
  • Intel oneAPI

Similar roles