Platform Engineer III
Cmegroup- Location
- Bangalore - Bagmane Tridib
- Workplace
- —
- Employment
- —
- Salary
- —
Posted 22d ago
Core Responsibilities
- 5+ Years experience, Implement, maintain, and optimize AI Platform capabilities deployed on Google Kubernetes Engine (GKE) in Google Cloud Platform (GCP), enabling product teams to build and deploy Generative AI applications and self-hosted models.
- Apply software and platform engineering best practices—including code quality, automated testing, CI/CD, GitOps, OTel telemetry, and documentation—to maintain reliable platform components and internal developer tooling.
- Collaborate directly with application, data science, and platform teams to collect feedback, troubleshoot integration issues, and implement cloud-native AI tools and observability standards.
- Stay current with advancements in AI infrastructure, GPU resource management, LLM serving engines, and cloud-native networking to continuously improve platform efficiency and developer experience.
- Build and integrate platform components, custom K8s resources, and telemetry pipelines using solid software design patterns within the broader GCP/GKE ecosystem.
- Contribute to team technical documentation, create reusable platform templates ("paved paths"), and assist other engineers in adopting platform standards.
- Execute on team technical roadmaps to improve platform reliability, lower operational overhead, and speed up delivery cycles for AI product features.
- Skills and Experience Requirements1. AI Infrastructure & Generative AI Experience
- Self-Hosted LLM Serving: Hands-on experience deploying, configuring, and scaling self-hosted Large Language Models (LLMs) on GKE using inference engines such as vLLM, NVIDIA NIM, or SGLang. Basic understanding of Kubernetes GPU allocation, multi-GPU nodes, and model execution requirements.
- Cloud-Native AI Tools: Experience working with or integrating emerging cloud-native AI infrastructure tools and control planes such as Agentgateway or Kagent to support agentic workflows and request routing.
- AI Agent & LLM Observability (OpenTelemetry): Practical experience configuring OpenTelemetry (OTel) instrumentation and collectors for AI applications. Familiarity with OTel GenAI Semantic Conventions (gen_ai.*) to capture traces, metrics, token usage, and latency across agentic execution flows, tool calls, and model calls.
- GenAI Frameworks: Hands-on experience integrating with GenAI frameworks (e.g., LangGraph, LangChain, Google Agent Development Kit/ADK) and cloud services (e.g., Google Vertex AI, Google Agentspace, Gemini APIs).
- Production Deployment: Experience deploying and supporting production workload pipelines in Kubernetes with attention to latency, availability, and resource utilization.
2. Google Cloud & Cloud-Native Networking
- GKE & GCP Knowledge: Solid, practical experience deploying and managing workloads on Google Kubernetes Engine (GKE), Google Compute Engine (GCE), and related GCP infrastructure.
- Service Mesh: Practical experience operating and troubleshooting Istio Ambient Mode (or sidecar-based Istio transitioning to Ambient) for mTLS, traffic routing, and service-to-service communication.
- Cloud-Native Tooling: Strong familiarity with container runtime environments, GPU device plugins/operators on Kubernetes, OpenTelemetry Collectors (OTLP), GitOps workflows (e.g., ArgoCD, Flux), Infrastructure as Code (Terraform), and standard security practices.
3. Engineering & Domain Standards
- Software Engineering Practices: Practical understanding of design patterns, unit/integration testing, clean code principles, and writing maintainable code.
- Programming Skills: Strong proficiency in Python or Go for writing automation, custom tooling, scripts, or operators.
- Teamwork & Communication: Strong collaboration skills with the ability to write clear documentation, work across team boundaries, and explain technical setup to fellow engineers.
CME Group: Where Futures are Made
CME Group is the world’s leading derivatives marketplace. But who we are goes deeper than that. Here, you can impact markets worldwide. Transform industries. And build a career by shaping tomorrow. We invest in your success and you own it – all while working alongside a team of leading experts who inspire you in ways big and small. Problem solvers, difference makers, trailblazers. Those are our people. And we’re looking for more.
At CME Group, we embrace our employees' unique experiences and skills to ensure that everyone’s perspectives are acknowledged and valued. As an equal-opportunity employer, we consider all potential employees without regard to any protected characteristic.
Important Notice
Recruitment fraud is on the rise, with scammers using misleading promises of job offers and interviews to solicit money and personal information from job seekers. CME Group adheres to established procedures designed to maintain trust, confidence and security throughout our recruitment process. Learn more here.
Skills
- Kubernetes
- GKE
- GCP
- Generative AI
- OpenTelemetry
- LLM
- vLLM
- NIM
- LangGraph
- LangChain
- Vertex AI
- Gemini
- Compute Engine
- Istio
- mTLS
- Argo CD
- Flux
- Terraform
- Python
- Go
More jobs at Cmegroup
All 22Network Ops Engineer
Cmegroup · Bangalore - Bagmane Tridib · yesterday
Mgr Pricing
Cmegroup · Chicago - 20 S. Wacker · USD 108,700–181,100/yr · 3d ago
Staff Data Reliability Engineer - India
Cmegroup · Bangalore - Bagmane Tridib · 8d ago
Director Data & Analytics API Experience - Data Services
Cmegroup · Belfast - Millennium House · 9d ago
Network Ops Engineer III- India
Cmegroup · Bangalore - Bagmane Tridib · 10d ago
Similar roles
Senior Backend Engineer
Zact · Mumbai, India · today
Principal AI Engineer - Hybrid in Bangalore
Smartsheet · Bangalore, INDIA · today
Forward Deployed Engineer
DevRev · India · today
Software Engineer
Vi · IN - Bengaluru, India · today
Cloud Developer
Hpe · Bengaluru, Karnātaka, India · today
Software Engineer - Packet Forwarding
Hpe · Bengaluru, Karnātaka, India · today