Senior Site Reliability Engineer
- Location
- Global
- Workplace
- Remote
- Employment
- Full Time
- Salary
- —
Posted 2mo ago
The employer’s full description could not be read from their board. This is a summary of the posting — follow the apply link for the original.
Responsibilities
- Design complex IT/OT architectures in cloud and on-prem
- Work directly with customers to understand environment and estimate effort
- Own customer solutions end-to-end from requirements to support
- Build or use reusable modules and bespoke solutions
- Deploy and manage Kubernetes-based infrastructure and stateful applications
- Participate in on-call rotation
- Own incidents through resolution and drive root cause analysis
- Build runbooks, alerts, and automation for incident prevention
- Provision and manage customer environments using Infrastructure-as-Code
- Implement and maintain GitOps workflows for in-cluster deployments
- Ensure infrastructure and application changes are declarative and version-controlled
- Automate self-healing and system updates
- Build and maintain monitoring, alerting, and dashboards
- Define SLIs and SLOs
- Contribute to operational standards, patterns, and processes
Requirements
- 5+ years in SRE, DevOps, or Infrastructure Engineering
- Strong Kubernetes skills in production environments
- Experience with GitOps tooling like ArgoCD, Rancher Fleet, or FluxCD
- Solid understanding of Infrastructure-as-Code concepts
- Experience with Terraform, Pulumi, or Crossplane
- Real incident response experience and on-call experience
- Ability to operate in ambiguity
- Clear communication skills for documentation and customer interaction
Preferred
- Azure experience
- Experience with SUSE ecosystem including SLE Micro, RKE2, Rancher, and Longhorn
- Industrial, manufacturing, or OT environment experience
- Familiarity with Inductive Automation's Ignition platform and MQTT
- Experience in a startup or small-team environment
Skills
- Kubernetes
- ArgoCD
- Rancher Fleet
- FluxCD
- Terraform
- Pulumi
- Crossplane
- Prometheus
- Loki
- Grafana
- Azure
- SLE Micro
- RKE2
- Rancher
- Longhorn
- Ignition
- MQTT
Similar roles
Senior Software Engineer
Coralogix · London, England, United Kingdom · today
Backend Engineer - Data Pipeline
Coralogix · Berlin, BE, Germany · today
Site Reliability Engineer (FedRAMP / Security)
Coralogix · New York, NY, United States · USD 170,000–350,000/yr · today
Sr Advanced Tech Product Owner
Honeywell · Bengaluru, Karnataka, India · today
Test Engineer
Aumovio · Chang Chun, Ji Lin, China · today
Senior Business Engineer - Ads
Reddit · Remote · USD 180,200–252,300/yr · today