Principal Site Reliability Engineer
Oracle- Location
- United Kingdom
- Workplace
- —
- Employment
- Full Time
- Salary
- —
Posted 17d ago
Candidates should have hands-on expertise with remote access technologies such as Ethernet, VPN, Load Balancing and BGP for secure and scalable route distribution. A strong understanding of Linux system processes, memory utilisation, disk and log management, network functionality, containerisation, and the TCP/IP stack is essential.
The role involves triaging and resolving Severity 1 and 2 incidents using logs, metrics, and CLI tools under pressure, including failed changes or system and process failures that directly impact customers in a 24/7 operational environment and potentially work directly with the customer.
This role requires the successful candidate to obtain UK Security Check (SC) clearance.
Key Responsibilities
-Work with the Virtual Networking team to share full-stack ownership of a collection of services and technology areas, providing operational support as part of an on-call rotation. Understand the end-to-end configuration, technical dependencies, and overall behavioural characteristics of production services. Take responsibility for the delivery of the mission-critical stack with a strong focus on security, resiliency, scalability, and performance.
-Hold authority for end-to-end performance and operability. Partner with global development teams to define and implement improvements in service architecture. -Clearly articulate the technical characteristics of services and technology areas, guiding development teams to engineer and deliver premier capabilities within the Oracle Cloud service portfolio.
-Develop and communicate a clear understanding of the scale, capacity, security, and performance attributes and requirements of the service and technology stack. -Demonstrate a solid grasp of automation and orchestration principles.
-Act as the ultimate escalation point for complex or critical issues that have not yet been documented as Standard Operating Procedures (SOPs). Apply a deep understanding of service topologies and their dependencies to troubleshoot issues and define mitigations. Understand and explain the impact of product architecture decisions on distributed systems
-Exhibit professional curiosity and a desire to develop a deep technical understanding of services and technologies.
-Ensure high quality, accurate and timely technical documentation of incidents, problems, changes, and standard operating procedures is maintained using tools such as Jira and Confluence
-Work is non-routine and highly complex, involving the application of advanced technical and business skills within the Virtual Networking specialisation of Oracle Cloud Infrastructure (OCI).
Qualifications
-Strong understanding of virtual network architecture, security, and automation
-Understanding of TCP/IP stack and routing concepts in Linux systems and networking environments. IPSEC, VPNs and BGP specifically.
-Experience with containerisation technologies and orchestration platforms.
-Solid understanding of Virtual Cloud Networks (VCNs) in public cloud environments.
-Experience with CI/CD systems and release automation tools.
-Experience in scripting languages such as Python or Shell.
-Familiarity with infrastructure automation tools such as Terraform and Chef.
-Possess leadership experience to ensure appropriate changes, upgrades, and enhancements are made based on the technical analysis.
-Must support network segmentation (e.g., security lists, network security groups, or firewalls).
-Deep Understanding of manipulating telemetry data (traffic flows, health status) using Grafana dashboards and MQL.
-Experience with major public cloud providers (e.g., Oracle Cloud Infrastructure OCI, or equivalent).
-Experience using Jira and Confluence for incident tracking, knowledge management, and ongoing technical documentation.
Core Responsibilities
Planning & Execution
- Manages and coordinates moderately complex tasks, monitoring timelines and deliverables to ensure timely completion and adherence to requirements for a moderately sized project or initiative. Efficiently delegates, monitors, and prioritizes work across multiple projects, providing technical oversight and adjusting plans to address shifts in resources or timelines.
Collaboration & Partnership
-Collaborates across the organization to align on expectations and achieve shared objectives. Leverages understanding of business leaders, stakeholders, and/or customers to ensure proposed solutions meet their needs. Supports inclusivity by actively seeking and listening to diverse perspectives, ensuring others feel heard and respected.
Problem Solving
- Identifies and addresses moderately complex issues by analyzing a wide range of data and/or information to identify solutions in accordance with standard practices. Proactively escalates unresolved or critical issues with a thorough assessment and suggests potential solutions. Reviews, contributes to, and documents problem solving strategies.
Continuous Learning
- Pursues learning opportunities to expand knowledge and skills and/or tools in new areas and stays abreast of the latest industry trends and best practices. Proactively seeks and leverages ongoing feedback and training to improve skills. Coaches and mentors junior team members, fostering continuous learning and knowledge sharing within and across teams.
Continuous Improvement
-Develops ideas, recommends updates, and/or collaborates on the implementation of process improvements to increase the efficiency and effectiveness of processes, protocols, and workflows across teams, and evaluates the impact on key stakeholders. Solicits feedback from others on ideas for alternative approaches and methods for continued improvement.
Performance and Development
-Contributes to the talent development pipeline by participating in candidate interviews, assessing candidates, and providing hiring recommendations.
-
Career Level - IC4
Skills
- VPN
- BGP
- Linux
- TCP/IP
- Oracle Cloud
- SOPS
- Jira
- Confluence
- OCI
- Python
- Shell
- Terraform
- Chef
- Grafana
More jobs at Oracle
All 313Principal Software Engineer, Core Infrastructure
Oracle · Nashville, TN, United States · today
Principal Platform Software Engineer - Robotics
Oracle · United States · Santa Clara, CA, United States · today
Senior Principal Linux Systems & RPM Development Engineer
Oracle · Santa Clara, CA, United States · Seattle, WA, United States · United States · USD 135,200–306,400/yr · today
Senior Platform Software Engineer
Oracle · Nashville, TN, United States · Austin, TX, United States · USD 92,500–209,500/yr · today
Principal Network Developer
Oracle · Austin, TX, United States · Nashville, TN, United States · United States · USD 102,300–209,500/yr · today
Similar roles
Security Engineer, Incident Response
Twilio · Ireland · United Kingdom · today
Security Engineer, Incident Response
Twilio · Ireland · United Kingdom · today
Product Security Engineer
Vercel · San Francisco, CA · New York, NY · London, England +1 · USD 208,000–312,000/yr · today
Senior Network Engineer
Jane Street · London, England, United Kingdom · today
Technical Support Engineer II - W&B EMEA
CoreWeave · London, England · today
Platform and Services Engineer - IAM (we have office locations in Cambridge, Leeds and London)
Genomics England · London, England, United Kingdom · today