Senior Site Reliability Engineer NEX
Candidate Experience site- Location
- Houston, TX, United States
- Workplace
- —
- Employment
- —
- Salary
- —
Posted 1mo ago
Reliability Engineering
- Design, implement, and operate scalable, resilient, and highly available systems on Google Cloud Platform.
- Improve service availability, latency, performance, scalability, and operational resilience.
- Define, implement, and track service-level indicators, service-level objectives, and error budgets.
- Perform capacity planning, performance analysis, and workload forecasting.
- Design and validate disaster recovery, backup, failover, and service-restoration capabilities.
- Implement and maintain secure cloud networking, IAM, workload identities, service accounts, and access-control practices.
- Partner with cybersecurity and identity teams to ensure infrastructure and services follow organizational security standards.
- Monitor cloud consumption and optimize resource utilization, performance, and cost efficiency.
- Identify operational risks and recommend improvements to cloud architecture and service design.
Automation and Platform Engineering
- Build and maintain cloud infrastructure using Terraform or comparable infrastructure-as-code tools.
- Automate repetitive operational activities and systematically identify, measure, and reduce manual toil.
- Build reusable infrastructure modules, deployment patterns, and operational tooling.
- Improve CI/CD pipelines to enable secure, repeatable, and reliable software delivery.
Observability and Incident Management
- Develop actionable alerts that identify meaningful service degradation while reducing alert fatigue and unnecessary operational noise.
- Create and maintain dashboards, runbooks, operational procedures, and troubleshooting documentation.
- Participate in a sustainable on-call rotation supporting production systems.
- Respond to production incidents, coordinate service restoration, and lead incident response when appropriate.
- Facilitate blameless postmortems and identify corrective and preventive actions.
- Use incident and operational data to improve system design, automation, monitoring, and response processes.
Collaboration and Service Ownership
- Partner with software engineering, data engineering, security, and product teams to improve application reliability and production operations.
- Promote shared responsibility for production reliability between application development and platform teams.
- Establish and document reliability standards, operational practices, and reusable engineering patterns.
- Provide technical guidance and coaching on SRE, cloud, Kubernetes, observability, and incident-management practices.
Required Knowledge, Skills, and Abilities
- Three or more years of experience in Site Reliability Engineering, platform engineering, DevOps, cloud engineering, production software engineering, or a similar role.
- Experience operating highly available systems in a 24/7 production environment.
- Hands-on experience operating workloads on Google Cloud Platform or another major public cloud platform.
- Strong experience managing compute, networking and data GCP services workloads
- Strong experience with containerization and orchestration technologies, including Docker and Kubernetes.
- Experience building and managing infrastructure with Terraform or a comparable infrastructure-as-code tool.
- Proficiency in Python, Go, Java, or another comparable programming language.
- Experience implementing or operating CI/CD pipelines using GitHub Actions,Azure DevOps, Bitbucket Pipelines, or comparable tools.
- Experience implementing observability using metrics, logs, traces, dashboards, and alerts.
- Experience participating in on-call rotations, responding to incidents, and contributing to postmortems.
- Understanding of SLIs, SLOs, error budgets, and other SRE principles.
- Ability to troubleshoot complex issues across application, infrastructure, network, data, and cloud-service layers.
- Ability to communicate effectively with engineering teams, business stakeholders, and operational personnel.
Minimum Qualifications
- Bachelor’s degree in Computer Science, Information Technology, Engineering, or a related field, or equivalent practical experience.
- 3+ years of experience in Site Reliability Engineering, platform engineering, cloud engineering, or DevOps.
- 3+ years of experience operating production workloads in GCP.
- Ability to understand and communicate in English at a level sufficient to issue, receive, and respond to safety-related and operations-related instructions.
Preferred Qualifications
- Google Cloud and/or Kubernetes certifications.
- Experience supporting data-intensive, streaming, analytics, or event-driven platforms.
- Experience establishing production-readiness, incident-management, or reliability-review processes.
- Experience working in the energy, oil and gas, industrial, IoT, field operations, or other operationally critical industries.
- Experience supporting technology environments that integrate cloud platforms with remote sites, field equipment, industrial systems, or edge computing.
Skills
- GCP
- IAM
- Terraform
- Kubernetes
- Docker
- Python
- Go
- Java
- GitHub Actions
- Azure DevOps
- Bitbucket
More jobs at Candidate Experience site
All 30Application Developer
Candidate Experience site · TX, United States · 10d ago
Digital Operations Engineer NEX
Candidate Experience site · Houston, TX, United States · 10d ago
Product Development Engineer- Power Systems
Candidate Experience site · Houston, TX, United States · 15d ago
Product Development Manager PDC
Candidate Experience site · Houston, TX, United States · 16d ago
Cement Lab Tech 3 NEX
Candidate Experience site · Midland, TX, United States · 16d ago
Similar roles
IT Support Specialist-III
Amperesand · Reno, Nevada, United States · today
Staff AI Engineering - Enterprise Architecture
American Express · Phoenix, AZ, United States · USD 144,250–256,250/yr · today
Field Service Technician II
Toshiba America Business Solutions · East Syracuse, NY, United States · today
Field Service Technician I
Toshiba America Business Solutions · Rochester, NY, United States · today
Field Service Technician II
Toshiba America Business Solutions · Buffalo, NY, United States · today
Incentives Analyst
MicroStrategy · Tysons Corner, VIRGINIA, United States · USD 66,400–119,600/yr · today