Inference Engineer – LLM & Speech AI
Soket.ai- Location
- Bengaluru, India
- Workplace
- Onsite
- Employment
- Full Time
- Salary
- —
Posted 4mo ago
We are looking for an experienced Inference Engineer to build and optimize high-performance inference systems for Large Language Models (LLMs), Speech AI systems, and multimodal AI workloads.
You will work on deploying production-grade
AI systems with a strong focus on
- low latency,
- high throughput,
- GPU efficiency,
- scalable serving infrastructure,
- distributed inference,
- and cost optimization.
This role sits at the intersection of
- systems engineering,
- deep learning infrastructure,
- distributed computing,
- and production AI deployment.
You will collaborate closely with
- ML researchers,
- platform engineers,
- speech AI teams,
- and product engineering teams.
Skills
- LLM
- Deep Learning
- Machine Learning
Similar roles
Senior Backend Engineer
Zact · Mumbai, India · today
Forward Deployed Engineer
DevRev · India · today
Cloud Developer
Hpe · Bengaluru, Karnātaka, India · today
Software Engineer - Packet Forwarding
Hpe · Bengaluru, Karnātaka, India · today
Systems Operations Senior Manager - Production Operations, AIOps, MLOps
Wf · Bengaluru, India · today
Manager, Software Engineering, ITC
N IKE, Inc . · Karnataka, India · today