JobHabor

Senior Site Reliability Engineer

Oracle
Location
BENGALURU, KARNATAKA, India
Workplace
Employment
Salary
Apply on the employer’s site

Posted 7d ago

Job Responsibilities

  • Improve the reliability, scalability, performance, and operational efficiency of assigned OCI Compute services and components.
  • Investigate and resolve complex production incidents; contribute to mitigation, recovery, RCA, and follow-up actions.
  • Own and improve service-level KPIs, SLOs, dashboards, alerting, deployment validation, and operational procedures for assigned systems.
  • Build automation and tooling to reduce operational toil and improve production safety.
  • Partner with development and infrastructure teams on service architecture, deployment, configuration, and reliability improvements.
  • Use observability, telemetry, event correlation, and AIOps capabilities to improve detection, diagnosis, and incident response.
  • Support upgrades, migrations, patching, capacity planning, performance tuning, security vulnerability management and production rollouts.
  • Troubleshoot distributed-system issues by analyzing service topology, dependencies, configuration, and failure modes.
  • Contribute to incident-management practices, operational readiness, and service ownership improvements.
  • Share technical knowledge and support team members through documentation, reviews, and collaboration.
  • Participate in a 12x7 on-call rotation and support response to customer-impacting incidents.

Mandatory Skills

  • 4–8 years of experience in SRE, Production Engineering, Cloud Operations, Systems Engineering, or a similar role.
  • Experience operating and improving highly available production systems.
  • Strong programming or scripting skills in Python, Java, Go, or similar languages.
  • Hands-on experience with Linux, cloud infrastructure, networking, compute, and storage.
  • Experience with production monitoring, alerting, dashboards, logs, metrics, and tracing.
  • Experience owning or improving service SLIs, SLOs, KPIs, and operational procedures.
  • Strong incident troubleshooting, RCA, debugging, and problem-solving skills.
  • Experience with deployment pipelines, release validation, automation, and change-management practices.
  • Understanding of distributed systems, service dependencies, capacity planning, and performance tuning.
  • Ability to work independently on technical problems and collaborate effectively with engineering teams.
  • Strong written and verbal communication skills.

Preferred Skills

  • Experience with OCI and cloud infrastructure services.
  • Experience with AIOps, anomaly detection, event correlation, predictive alerting, or automated remediation.
  • Experience with Kubernetes, containers, infrastructure-as-code, and CI/CD.
  • Experience with service migrations, fleet maintenance, upgrades, patching, or production rollouts.
  • Experience with architecture reviews, operational-readiness reviews, and post-incident improvements.
  • Experience contributing to technical initiatives, knowledge sharing, code reviews, or operational improvements within the team.
  • Familiarity with security, compliance, and access-control practices in production environments.

Self-Test Questions

  • Do you have 4–8 years of relevant SRE, Production Engineering, Cloud Operations, or Systems Engineering experience?
  • Have you independently operated or improved a production service, system, or infrastructure component?
  • Can you investigate production incidents and contribute to mitigation, recovery, RCA, and follow-up actions?
  • Do you have hands-on experience with Linux, cloud infrastructure, distributed systems, networking, compute, or storage?
  • Are you proficient in Python, Java, Go, or a similar language for automation, tooling, and troubleshooting?
  • Have you built or improved automation, deployment validation, CI/CD pipelines, or operational tooling?
  • Do you have experience with monitoring, alerting, logs, metrics, tracing, and service health indicators such as SLOs or KPIs?
  • Can you work independently on assigned technical problems, collaborate with partner teams, and participate in a 12x7 on-call rotation?

Career Level - IC3

Skills

  • OCI
  • Python
  • Java
  • Go
  • Linux
  • Kubernetes

More jobs at Oracle

All 297

Similar roles