JobHabor

Untitled role

Io Tech Solutions
Location
Location not stated
Workplace
Employment
Salary
Apply on the employer’s site

Senior Engineer – AI Platform IO TECH SOLUTIONS LIMITED | Career Page

Senior Engineer – AI Platform

Job Description

Role Overview We are looking for a hands-on Senior Engineer to design, build, and operate an enterprise AI platform for a digital bank. Our goal is to provide a fully governed, safe, and efficient AI environment that staff actually want to use.

You will take ownership across three core areas

an enterprise-grade inference gateway, identity and distribution systems, and self-hosted GPU inference clusters.

Core Responsibilities

Inference Serving & On-Premise GPU Operations

  • Size hardware capacity and collaborate with infrastructure teams to handle hypervisors, rack/power setups, and GPU passthrough across bare-metal or data center environments.
  • Deploy, operate, and fine-tune model serving frameworks using techniques like continuous batching, quantization, tensor parallelism, and KV cache optimization.
  • Implement tenancy admission controls, quota management, queuing, and failover mechanisms to cloud inference when local capacity is saturated.
  • Track and publish key performance indicators—such as p95 time-to-first-token, tokens per second, and cost per million tokens—to evaluate fixed OpEx break-even points against cloud alternatives.

Governed Gateway Architecture & Identity Integration

  • Construct and maintain an inference gateway that handles rate limiting, per-user key management, model allowlists, and strict zero data retention policies.
  • Connect governance systems directly to corporate single sign-on (OIDC / OAuth2) and automate onboarding, offboarding, and role claims via SCIM.
  • Ensure total financial transparency by attributing real-time API spend, token usage, and operational costs down to individual users and departments.

Tooling Distribution & User Enablement

  • Package, distribute, and manage agent configurations and execution harnesses across a diverse enterprise machine fleet.
  • Maintain a secure, CI-driven registry for versioned skills, knowledge-base bundles, and agent modules featuring seamless rollback capabilities.
  • Onboard, train, and support non-technical teams (such as risk, finance, and operations) to maximize productivity using approved AI tools.

What We Are Looking For

Essential Qualifications & Skills

  • Production Platform Engineering: 5+ years of experience managing critical production infrastructure, with direct involvement in incident response and on-call rotations.
  • Core Technical Stack: Proficiency in Python (primary) and TypeScript, alongside hands-on experience with Kubernetes, Linux, containers, Infrastructure as Code, and CI/CD pipelines.
  • Enterprise Identity Systems: Strong grasp of JWT claims, OIDC, OAuth2, SCIM protocols, and enterprise directory integrations.
  • Gateway & Multi-Tenancy Architecture: Background in building or maintaining API/LLM gateways, rate limits, quota systems, usage metering, and multi-tenant isolation.
  • On-Premise Infrastructure Comfort: Experience navigating capacity planning, hypervisors, hardware passthrough, and enterprise change management processes.
  • FinOps & Attribution Mindset: Ability to evaluate platforms through unit economics, capacity utilization, and cost attribution.
  • Technical Enablement: Skill in simplifying complex AI tools and training non-technical colleagues to use them effectively.

Preferred / Bonus Qualifications

  • Experience with Model Context Protocol (MCP), sovereign hosting constraints, model evaluation, or highly regulated financial services environments.
  • Hands-on expertise tuning high-throughput GPU model serving stacks such as vLLM, SGLang, Triton, or TGI.

Platform Guiding Principles

  • Usability First: Approved pathways must offer a better user experience than unapproved workarounds to prevent control bypasses.
  • Attribution Before Scale: Total visibility into spend and usage per user is required before expanding system scale.
  • Allocated Capacity: GPU resources are scarce and access is governed through explicit capacity allocation policies.
  • Pragmatic Ownership: Open-source or licensed solutions are used for standard components, while internal focus is dedicated to security and governance boundaries.

Required Skills

JWT Environment Hardware Data Center Incident Response Performance CI/CD pipelines Data Cloud API Support Access AI Usability Transparency Financial Services Pipelines Operations Ownership User Experience Onboarding CI/CD Components Architecture Change Management Optimization Infrastructure Economics Integration Kubernetes TypeScript Security Linux Finance Planning Design Engineering Training Python Management

Apply for this job

Or refer someone

Share

  • Facebook
  • Line
  • LinkedIn
  • X (Formerly Twitter)
  • Whatsapp
  • Email

Skills

  • OpenID Connect
  • OAuth 2.0
  • SCIM
  • Python
  • TypeScript
  • Kubernetes
  • Linux
  • JWT
  • LLM
  • Model Context Protocol
  • vLLM
  • Triton

More jobs at Io Tech Solutions

All 78