Principal Site Reliability Engineer, Infrastructure Observability
Troweprice- Location
- London, Warwick Court
- Workplace
- —
- Employment
- —
- Salary
- —
Posted 28d ago
KM5
Role Summary
In this role as Principal Site Reliability Engineer, Infrastructure Observability you will help formulate, develop, and implement a team of Site Reliability Engineers (SREs) focused on the observability, sustainability, scalability, measurability and recoverability of T. Rowe Price’s innovative cloud & on-prem solutions by leveraging automation and best-of-breed tools. The successful candidate will have a strong operations & engineering background, is hands-on when needed, and has expertise in the cloud environments (public, private), infrastructure operations, DevOps practices, CI/CD toolchain and systems, code build and deployment, incident response, and 24x7 monitoring and support.
The candidate will also have extensive experience operating within a SRE function within a complex, distributed environment. They will have a demonstrated ability to work horizontally and vertically within an organization with diverse partners and sponsor groups.
Responsibilities
- Possesses extensive knowledge in own area of expertise and extensive in-depth knowledge of the broader portfolio for comprehensive understanding of up/downstream impacts across technology infrastructure
- Responsibility for the design of technology solutions to prevent or minimize service disruptions
- Prevents technology service disruptions through technology solution recommendations and automations
- Fosters a culture of deep learning through blameless post-mortems to improve the shared goal of reliability across services
- Transform operations teams by facilitating internal change to adopt SRE standard methodologies across the organization and driving strategic growth in this area within Global Technology
- Analyzes incidents impacting technology availability for high-level trends across the broad portfolio
- Drive initiatives to reduce or prevent technology failures in a complex, distributed technology environment
- Pulls together information from disconnected systems into cohesive views of the technology portfolio for identifying trends, redundancies, and risk
- Demonstrates outstanding awareness of the complexities of the tech and asset management industries
- May lead initiatives of varying degrees of complexity that span multi-functional areas and of varying degrees of complexity
- Contributes to definition of target state architecture and design of the technology environment
Qualifications
Required
- Bachelor's degree or the equivalent combination of education and relevant experience AND 10+ years of experience designing and operating cloud infrastructure with senior‑level impact.
- 5+ years building and supporting solutions in Amazon AWS
- 5+ years of experience building and running a DevOps and/or SRE function
- Experience with implementation and operation of the chaos model at scale
- Strategic and program-level implementation experience
- Demonstrable experience implementing new technology, tools, and platforms
- System administration and scripting experience
- Demonstrable experience leveraging automation to proactively prevent or quickly remediate incidents
- Fluent in multiple programming languages (e.g., Python, Java, GO, Node.js, .Net Core, etc)
- Proficiency with database development (SQL Server, PostgreSQL, MySQL, etc)
- Proficiency with defining, right-sizing, tracking, and reporting on Service Level Objectives (SLOs), Service Level Indicators (SLIs), system availability, and the progress and outcomes related to reliability
- Experience with implementing and managing Error Budgets
- Proficiency with understanding and explaining incident situations and their recovery plans to prevent recurrence
- Knowledge/experience driving dashboard standardization across the ecosystem for observability, APM and infrastructure monitoring, and application-specific logging
- Knowledge/experience with observability tools such as New Relic, SolarWinds DPA, Elastic Stack, Prometheus, Grafana, Splunk, and cloud native tools
- Knowledge/experience with cloud management tools such as Ansible, Terraform, Vault, and Vagrant
- Works independently, with guidance in only the most complex situations
- Makes sound decisions with limited facts or resources
- Balances strategic and pragmatic concerns when solving problems
- Adjusts communication style and materials to suit a given audience
- Able to clearly articulate operational principles, practices, and policies
- Stays abreast of industry trends and technologies
- Accountable for work of self and others; sets standards around which others will operate
- Maintains a broad internal professional network and knows when to engage/activate it
- Develops or mentor’s diverse talent on the team
- Ability to be on-call and/or work during off-hours
Preferred
- Cloud or SRE‑related certifications
- Working knowledge of Azure
Work Flexibility
This role is eligible for hybrid work, with up to three days per week from home.
Skills
- Deep Learning
- AWS
- Python
- Java
- Go
- Node.js
- .NET Core
- SQL Server
- PostgreSQL
- MySQL
- New Relic
- ELK Stack
- Prometheus
- Grafana
- Splunk
- Ansible
- Terraform
- HashiCorp Vault
- Vagrant
- Azure
More jobs at Troweprice
All 9Senior AI Software Engineer - TRP Labs London
Troweprice · London, Warwick Court · 4d ago
Senior Analyst, Data Analytics
Troweprice · London, Warwick Court · 11d ago
Senior Software Engineer - ESG Techology
Troweprice · London, Warwick Court · 23d ago
Lead Splunk Administrator
Troweprice · London, Warwick Court · 28d ago
Lead AI Software Engineer - TRP Labs London
Troweprice · London, Warwick Court · 2mo ago
Similar roles
Senior Network Engineer
Jane Street · London, England, United Kingdom · today
Technical Support Engineer II - W&B EMEA
CoreWeave · London, England · today
Platform and Services Engineer - IAM (we have office locations in Cambridge, Leeds and London)
Genomics England · London, England, United Kingdom · today
Technical Support (Telecoms) - 2nd Line
Radius · Crewe, England, United Kingdom · today
AI / Machine Learning Engineer Leader
Prolific Academic · UK · USD 200/hr · today
Technical Program Manager - International
MS Career - Contingent Worker · United Kingdom · today