Senior Staff Site Reliability Engineer – Compute Platform
NVIDIA- Location
- India, Bengaluru
- Workplace
- —
- Employment
- —
- Salary
- —
Posted yesterday
NVIDIA is seeking a Senior Staff SRE to build and operate reliable, scalable compute platforms that support global engineering workloads. This role spans Kubernetes, KubeVirt, bare-metal infrastructure, automation, observability, and AI-enabled operations. Join a team that solves complex infrastructure challenges, builds durable automation, and improves the reliability and operational experience of critical compute services.
What you’ll be doing
- Build, operate, and improve large-scale Kubernetes, KubeVirt, Linux, container, and bare-metal compute platforms, with a focus on performance, capacity, reliability, and operational scale.
- Lead bare-metal provisioning and lifecycle management in data centers, including PXE boot, DHCP, DNS, OS provisioning, hardware validation, and fleet automation.
- Develop automation, self-service capabilities, and observability solutions using APIs, Python or Go, Infrastructure as Code, configuration management, metrics, logs, traces, and service-health data.
- Define and operate SLOs, SLIs, error budgets, alerting, and incident-response practices; lead complex incident investigations, corrective actions, and blameless postmortems.
- Partner with infrastructure, security, hardware, data-center, and application teams to deliver global platform initiatives, and participate in an on-call rotation.
What we need to see
- BS in Computer Science, Engineering, a related technical field, or equivalent experience, plus 10+ years operating production infrastructure or platform services.
- Strong expertise in Kubernetes administration, KubeVirt, Docker, containerization, microservices, Linux systems, and resolving distributed-system challenges.
- Experience deploying and operating bare-metal infrastructure in a data-center environment, including provisioning, networking, operating-system lifecycle management, and hardware automation.
- Proficiency in Python, Go, or a comparable programming language, with experience building RESTful services and integrating infrastructure APIs.
- Experience with Infrastructure as Code and automation tools such as Terraform, Ansible, Chef, or Puppet, along with a solid understanding of TCP/IP networking and infrastructure security.
- Strong SRE and observability experience, including SLIs, SLOs, error budgets, incident management, monitoring, logging, tracing, and tools such as OpenTelemetry, Prometheus, Grafana, ELK Stack, or Splunk.
- Clear written and interpersonal communication skills, with a record of delivering practical, scalable solutions to complex technical problems.
Ways to stand out from the crowd
- Experience operating HPC, AI, GPU-accelerated, or general-purpose bare-metal compute infrastructure, including GPU-enabled Kubernetes or KubeVirt clusters.
- Expertise with VMware vSphere, Red Hat OpenShift, KVM, Firecracker, OpenStack, or Nutanix AHV.
- Experience applying generative AI or agentic workflows to improve infrastructure diagnostics, reduce operational toil, and accelerate incident resolution.
- Experience building secure, integrated operational platforms using APIs, RBAC, service accounts, secrets management, audit controls, workflow orchestration, and infrastructure or incident-management systems.
- Demonstrated delivery of complex, high-impact infrastructure projects.
Skills
- Kubernetes
- Linux
- DHCP
- DNS
- Python
- Go
- Docker
- Terraform
- Ansible
- Chef
- Puppet
- TCP/IP
- OpenTelemetry
- Prometheus
- Grafana
- ELK Stack
- Splunk
- VMware
- OpenShift
- Generative AI
- RBAC
More jobs at NVIDIA
All 432Silicon Performance, Power and Binning Tools Engineer
NVIDIA · China, Shanghai · today
Pre Production Engineer
NVIDIA · Taiwan, Taipei · today
IC Test Program Validation and Verification Engineer
NVIDIA · Israel, Yokneam · today
Senior AI STA Engineer, Sub-chip
NVIDIA · Israel, Yokneam · Israel, Beer Sheva · Israel, Tel Aviv · today
Senior Systems Software Engineer
NVIDIA · India, Pune · India, Hyderabad · India, Bengaluru · yesterday
Similar roles
Cloud DevOps Engineer - Data QA
Southwest Airlines · India Office · today
DevOps Engineer
Ensono · Bengaluru, India · Chennai, India · Hyderabad, India +1 · today
DevOps Engineer
Sonicwall · Bengaluru, Karnataka, India · today
Software Dev Senior Engineer – SRE & Cloud Reliability
Sonicwall · Pune, Maharashtra, India · today
Senior DevSecOps Engineer - AWS, EKS, GitLab CI/CD
Sonicwall · Pune, Maharashtra, India · today
Software Engineering Associate (Java)
DTCC Candidate Experience Site · Hyderabad, India · Chennai, India · today