JobHabor

Site Reliability Engineer II

American Express
Location
Bengaluru, KA, India · Chennai, TN, India
Workplace
Hybrid
Employment
Full Time
Salary
Apply on the employer’s site

Posted yesterday

Site Reliability Engineer II collaborates with engineering teams to enhance system resilience, scalability, and performance through feature development, automation, architectural design, resiliency testing, and disaster recovery planning, while promoting best practices for continuous improvement.

  • Collaborates with Software Engineering teams to design, develop, and implement features that enhance system resilience, scalability, and performance, while identifying and addressing potential system bottlenecks and failure points with guidance from senior colleagues
  • Develops and implements automation tools and frameworks, including infrastructure as code (IaC) practices to streamline operational workflows, deployment processes, and infrastructure management, with guidance from peers and leaders
  • Collaborates with senior engineers to contribute to the architectural design of systems, ensuring that reliability, scalability, and performance considerations are integrated into design discussions and decision-making processes
  • Collaborates in the design and execution of chaos engineering experiments and other resiliency testing, analyzing results and implementing improvements to enhance system robustness and recovery capabilities, with guidance from peers and leaders
  • Develops and implements of disaster recovery plans and business continuity strategies, ensuring systems can recover quickly and effectively from unexpected disruptions
  • Collaborates with seniors to promote and implement best practices such as error budgeting, service-level objectives (SLOs), and service-level indicators (SLIs), contributing to a culture of continuous improvement and reliability
  • Collaborates and co-creates effectively with teams in product and the business to align technology initiatives with business objectives

Education Qualifications

  • Bachelor’s degree in Computer Science, Information Technology, Engineering, and/or comparable experience; advance degree preferred
  • Knowledge of modern observability stack – Splunk, Elastic Search, Prometheus, Grafana
  • Knowledge of containerization technologies (e.g., Kubernetes, Docker) and microservices architecture
  • Knowledge of observability tools and methodologies, including experience with logging, monitoring, tracing, and performance analysis platforms
  • Knowledge of cloud-based Site Reliability Engineering (SRE) practices and experience with public cloud platforms such as AWS, Azure, or Google Cloud

Work Experience

  • Experience in software development, or technology operations, with a focus on Site Reliability Engineering
  • Experience in Linux/Unix systems, object-oriented programming languages (e.g., Java), scripting languages (e.g., Python, Bash), and cloud platforms (e.g., AWS, Azure, GCP)

Licenses and Certifications

  • Advanced certification in Site Reliability Engineering (SRE) or related is a plus

Skills

  • Splunk
  • Elastic
  • Prometheus
  • Grafana
  • Kubernetes
  • Docker
  • AWS
  • Azure
  • GCP
  • Linux
  • Unix
  • Java
  • Python
  • Bash

More jobs at American Express

All 244

Similar roles