Senior Site Reliability Engineer
Oracle- Location
- BENGALURU, KARNATAKA, India
- Workplace
- —
- Employment
- Full Time
- Salary
- —
Posted 8d ago
Job Responsibilities
- Improve the reliability, scalability, performance, and operational efficiency of assigned OCI Compute services and components.
- Investigate and resolve complex production incidents; contribute to mitigation, recovery, RCA, and follow-up actions.
- Own and improve service-level KPIs, SLOs, dashboards, alerting, deployment validation, and operational procedures for assigned systems.
- Build automation and tooling to reduce operational toil and improve production safety.
- Partner with development and infrastructure teams on service architecture, deployment, configuration, and reliability improvements.
- Use observability, telemetry, event correlation, and AIOps capabilities to improve detection, diagnosis, and incident response.
- Support upgrades, migrations, patching, capacity planning, performance tuning, security vulnerability management and production rollouts.
- Troubleshoot distributed-system issues by analyzing service topology, dependencies, configuration, and failure modes.
- Contribute to incident-management practices, operational readiness, and service ownership improvements.
- Share technical knowledge and support team members through documentation, reviews, and collaboration.
- Participate in a 12x7 on-call rotation and support response to customer-impacting incidents.
Mandatory Skills
- 4–8 years of experience in SRE, Production Engineering, Cloud Operations, Systems Engineering, or a similar role.
- Experience operating and improving highly available production systems.
- Strong programming or scripting skills in Python, Java, Go, or similar languages.
- Hands-on experience with Linux, cloud infrastructure, networking, compute, and storage.
- Experience with production monitoring, alerting, dashboards, logs, metrics, and tracing.
- Experience owning or improving service SLIs, SLOs, KPIs, and operational procedures.
- Strong incident troubleshooting, RCA, debugging, and problem-solving skills.
- Experience with deployment pipelines, release validation, automation, and change-management practices.
- Understanding of distributed systems, service dependencies, capacity planning, and performance tuning.
- Ability to work independently on technical problems and collaborate effectively with engineering teams.
- Strong written and verbal communication skills.
Preferred Skills
- Experience with OCI and cloud infrastructure services.
- Experience with AIOps, anomaly detection, event correlation, predictive alerting, or automated remediation.
- Experience with Kubernetes, containers, infrastructure-as-code, and CI/CD.
- Experience with service migrations, fleet maintenance, upgrades, patching, or production rollouts.
- Experience with architecture reviews, operational-readiness reviews, and post-incident improvements.
- Experience contributing to technical initiatives, knowledge sharing, code reviews, or operational improvements within the team.
- Familiarity with security, compliance, and access-control practices in production environments.
Self-Test Questions
- Do you have 4–8 years of relevant SRE, Production Engineering, Cloud Operations, or Systems Engineering experience?
- Have you independently operated or improved a production service, system, or infrastructure component?
- Can you investigate production incidents and contribute to mitigation, recovery, RCA, and follow-up actions?
- Do you have hands-on experience with Linux, cloud infrastructure, distributed systems, networking, compute, or storage?
- Are you proficient in Python, Java, Go, or a similar language for automation, tooling, and troubleshooting?
- Have you built or improved automation, deployment validation, CI/CD pipelines, or operational tooling?
- Do you have experience with monitoring, alerting, logs, metrics, tracing, and service health indicators such as SLOs or KPIs?
- Can you work independently on assigned technical problems, collaborate with partner teams, and participate in a 12x7 on-call rotation?
Career Level - IC3
Skills
- OCI
- Python
- Java
- Go
- Linux
- Kubernetes
More jobs at Oracle
All 313Principal Software Engineer, Core Infrastructure
Oracle · Nashville, TN, United States · today
Principal Platform Software Engineer - Robotics
Oracle · United States · Santa Clara, CA, United States · today
Senior Principal Linux Systems & RPM Development Engineer
Oracle · Santa Clara, CA, United States · Seattle, WA, United States · United States · USD 135,200–306,400/yr · today
Senior Platform Software Engineer
Oracle · Nashville, TN, United States · Austin, TX, United States · USD 92,500–209,500/yr · today
Principal Network Developer
Oracle · Austin, TX, United States · Nashville, TN, United States · United States · USD 102,300–209,500/yr · today
Similar roles
Analyst IT – Quality
Mattel · Hyderabad, India · today
Principal Network Engineer
Hewlett Packard Enterprise · Bangalore, Karnataka, India · today
Principal Network Engineer
Hpe · Bangalore, Karnataka, India · today
Cloud DevOps Engineer - Data QA
Southwest Airlines · India Office · today
DevOps Engineer
Ensono · Bengaluru, India · Chennai, India · Hyderabad, India +1 · today
DevOps Engineer
Sonicwall · Bengaluru, Karnataka, India · today