SRE Lead
Chubb External- Location
- Malaysia
- Workplace
- —
- Employment
- Full Time
- Salary
- —
Posted 2mo ago
Key Objective
- Lead the Site Reliability Engineering function to define and drive the organisation’s reliability engineering strategy — bridging software development and operations through engineering discipline, not manual process.
- Own the end-to-end reliability posture of production systems: define SLO/SLI frameworks, govern error budgets, and enforce production-readiness standards to protect business continuity.
- Build, mentor, and scale a high-performing SRE team that prioritises engineering over toil — automating manual work, embedding reliability into the SDLC, and driving down mean time to recovery through systematic improvement.
- Champion observability-led engineering through full-stack Dynatrace adoption, AIOps integration, and data-driven reliability decision-making at every layer of the stack.
- Serve as the primary reliability engineering partner to development and platform leadership, shaping architecture decisions, release policies, and automation strategy.
Key Responsibilities
- Define and drive the SRE strategy and multi-year roadmap aligned to business priorities.
- Lead and develop the SRE team, including hiring, onboarding, performance management, career development, and succession planning.
- Own incident management, including severity classification, escalation, response SLAs, and leadership of major incidents.
- Champion blameless postmortems, root cause analysis, and implementation of systemic fixes.
- Establish and govern SLOs, SLIs, and error budgets, ensuring reliability targets are aligned to business needs.
- Drive resilience engineering, including chaos engineering, GameDays, production readiness reviews, and failure mode analysis.
- Reduce toil through automation, improved runbooks, and continuous operational improvement.
- Own observability and alerting standards, including monitoring strategy, dashboards, and alert quality.
- Partner with engineering, architecture, product, and leadership teams to embed reliability into design and delivery.
- Represent the SRE function in senior forums and provide reporting on reliability, risk, and operational performance.
Qualifications
- Degree in Computer Science, Software Engineering, IT, or a related technical field.
- 8+ years’ experience in software engineering, platform reliability, or SRE, including 3+ years in people leadership.
- Strong hands-on coding ability in Python, Go, or similar, with experience building automation and self-healing solutions.
- Proven experience leading enterprise-scale SRE or platform reliability functions.
- Experience defining and operating RTO/RPO and SLI/SLO frameworks.
- Strong background in observability, production readiness, error budgets, and chaos engineering.
- Experience leading on-call models, incident response, and executive stakeholder engagement.
- Solid understanding of SDLC, Agile, and DevOps delivery models.
- ITIL Foundation is desirable.
Managerial & Soft Skills
- Proven people leader with experience building and coaching high-performing teams.
- Strategic thinker who can turn business priorities into reliability roadmaps.
- Strong communicator who can explain technical risk in business terms.
- Calm and decisive during major incidents.
- Influential partner across engineering, product, and leadership teams.
- Strong advocate for developer experience and sustainable on-call practices.
- Data-driven and able to balance reliability, speed, and cost.
- Champions psychological safety, continuous learning, and operational excellence.
Technical Skills
- Expert in Dynatrace, with experience in observability, monitoring, SLOs, tracing, and log management.
- Proficient in Grafana, Prometheus, Splunk, ELK, Azure Monitor, and Log Analytics.
- Strong knowledge of OpenTelemetry and telemetry pipeline design.
- Experience with ServiceNow, CI/CD tools, Kubernetes, Docker, Terraform, and Bicep.
- Familiar with Java, .NET, databases, APIs, Kafka, and cloud platforms, especially Azure.
- Experience with AIOps, AI-assisted triage, and automation tooling.
- Able to support reliability engineering through scripting, auto-remediation, and operational automation.
Desired
- Experience in insurance or financial services.
- Dynatrace, Azure, ITIL 4, or Google Cloud/SRE-related certifications.
- Experience with chaos engineering, AIOps, MLOps, and FinOps.
- Strong analytical skills and experience working with large operational datasets.
Skills
- Dynatrace
- Python
- Go
- Grafana
- Prometheus
- Splunk
- ELK Stack
- Azure Monitor
- Azure Log Analytics
- OpenTelemetry
- ServiceNow
- Kubernetes
- Docker
- Terraform
- Bicep
- Java
- .NET
- Kafka
- Azure
- GCP
- MLOps
More jobs at Chubb External
All 146Software Engineer
Chubb External · Telangana, India · yesterday
AI Engineer (Fluent in Mandarin & English)
Chubb External · London, United Kingdom · yesterday
Lead Data Engineer - Quantexa
Chubb External · Telangana, India · 2d ago
Chubb Life: Supervisor - Partnership Operations Development
Chubb External · Bangkok, Thailand · 3d ago
Senior AI/ML Engineer II
Chubb External · Bangalore, Karnataka, India · 3d ago
Similar roles
IT Support Engineer
Enovix · Penang, MY · today
Hardware NPI / Debug Engineer - Malaysia
Arista Networks · George Town, Penang, Malaysia · today
(Senior) Product Analyst
Delivery Hero · Kuala Lumpur, Malaysia · today
Technician 2, Information Technology
Western Digital · Bayan Lepas, Penang, Malaysia · yesterday
AI & Automation Intern - KL
Teleperformance · Malaysia · yesterday
AI & Automation Intern - Penang
Teleperformance · Malaysia · yesterday