JobHabor

Manager, Site Reliability Engineering

Aya Healthcare

Location
US
Workplace
Remote
Employment
Full Time
Salary
USD 230,000–255,000/yr
Apply on the employer’s site

Posted 3mo ago

The employer’s full description could not be read from their board. This is a summary of the posting — follow the apply link for the original.

Responsibilities

  • Lead and grow the SRE team
  • Set the operating cadence for the team
  • Build a culture of blameless learning
  • Partner closely with DevSecOps, Security Engineering, DRE, Incident & Change Management, and product engineering leadership
  • Own the reliability strategy for customer-facing products and internal platforms
  • Define SLOs, SLIs, and error budgets
  • Lead major incident response as senior incident commander
  • Champion proactive reliability
  • Manage software release support and 24/7 on-call escalation rotations
  • Build the AIOps practice
  • Operationalize AI-assisted workflows
  • Pilot and scale agentic remediation
  • Evolve the observability platform
  • Drive platform unit economics
  • Communicate outcomes to executive, product, and customer-facing stakeholders
  • Uphold HIPAA, PHI, and security obligations

Requirements

  • 10+ years in SRE, DevOps, Platform Engineering, or related production-operations roles
  • 4+ years of direct people management experience
  • Demonstrated ownership of reliability outcomes for customer-facing SaaS
  • Deep Azure experience (3+ years operating production workloads on Azure)
  • Hands-on depth in AKS, networking, identity, and platform services
  • Modern observability fluency (Datadog or equivalent)
  • AI in operations experience
  • Incident command experience
  • Regulated-environment instinct (HIPAA, PHI, SOC 2)
  • Executive-grade communication
  • Bachelor's degree in Computer Science, Information Technology, Engineering, or related field, or equivalent experience

Preferred

  • Cloudflare at the edge
  • IaC at scale (Terragrunt and Terraform)
  • CI/CD maturity (GitHub Actions)
  • Container platform depth (Kubernetes/AKS)
  • ITSM integration (ServiceNow)
  • Identity ecosystem (Okta / Entra ID / M365)
  • Chaos and resilience engineering
  • FinOps fluency
  • Agile delivery (Scrum/Kanban)

Skills

  • Azure
  • AKS
  • Datadog
  • Cloudflare
  • Terragrunt
  • Terraform
  • GitHub Actions
  • Kubernetes
  • ServiceNow
  • Okta
  • Entra ID
  • M365

Similar roles