JobHabor

Senior Site Reliability Engineer

Veeam

Location
United States
Workplace
Remote
Employment
Full Time
Salary
USD 125,300–320,900/yr
Apply on the employer’s site

Posted 2mo ago

The employer’s full description could not be read from their board. This is a summary of the posting — follow the apply link for the original.

Responsibilities

  • Get up to speed on the full platform
  • Write and maintain runbooks, architecture docs, and operational guides
  • Design infrastructure for high availability and fault tolerance on Azure
  • Define SLIs, SLOs, and error budgets
  • Run incident response and blameless postmortems
  • Identify reliability risks and build practical remediation plans
  • Close observability gaps
  • Set alerting, telemetry, and monitoring standards
  • Build automation to reduce toil and support fleet management
  • Work with IaC, CI/CD, deployment automation, and config management
  • Build and maintain testing, canary deployment, and release validation pipelines
  • Integrate chaos engineering and monitoring tools
  • Work across product, platform, security, legal, compliance, and operations teams
  • Own problems end-to-end
  • Mentor other engineers and help spread SRE practices

Requirements

  • 7+ years in Software Engineering, with 3+ years in SRE, Platform Engineering, or similar
  • Experience with Government or Sovereign Cloud (e.g., Azure Government, AWS GovCloud)
  • Experience in regulated compliance environments — government (FedRAMP, CMMC, IL2/IL4/IL5), financial (PCI-DSS, SOX), or healthcare (HIPAA, HITRUST)
  • Understand how compliance shapes architecture and operations
  • Strong experience building and running production services on cloud infrastructure (Azure preferred, including Azure Government)
  • Able to learn large, complex platforms quickly with limited guidance
  • Comfortable building understanding from code, docs, and architecture artifacts when direct environment access is restricted
  • Can investigate systems independently and produce clear docs, risk assessments, and improvement plans
  • Comfortable working across teams — engineering, product, security, compliance, operations
  • Programming skills in one or more of: TypeScript/JS, Go, Java, C#, or similar
  • Experience with monitoring and observability tools (e.g., Prometheus, Grafana, OpenTelemetry, ELK stack)
  • Experience with IaC (Terraform, Terragrunt, Pulumi) and container orchestration (Kubernetes)
  • Experience with CI/CD and GitOps tooling — GitHub Actions, Azure DevOps, GitLab CI, ArgoCD, FluxCD, or Dagger
  • Solid grasp of distributed systems, networking, and cloud-native architecture
  • Clear written and verbal communication skills

Preferred

  • Experience on B2B SaaS platforms in regulated or government markets
  • Background in chaos engineering, resilience testing, or performance/load testing
  • Have built an SRE or reliability function from scratch before
  • Experience across mixed environments — modern cloud-native and older legacy systems
  • Familiar with AI-first development workflows — using LLM-powered tools for infrastructure automation, code generation, and documentation

Skills

  • Microsoft TFS
  • Azure DevOps
  • Git
  • BitBucket
  • Azure
  • Entra ID
  • API Management
  • Cosmos DB
  • Storage services
  • Azure Functions
  • static website hosting
  • Azure security
  • Azure ARM templates
  • AWS CloudFormation
  • Terraform
  • Serverless Framework
  • Azure Monitor
  • AppInsights
  • Elastic Stack
  • TypeScript
  • JavaScript
  • Go
  • Java
  • C#
  • Prometheus
  • Grafana
  • OpenTelemetry
  • ELK stack
  • Terragrunt
  • Pulumi
  • Kubernetes
  • GitHub Actions
  • GitLab CI
  • ArgoCD
  • FluxCD
  • Dagger

Similar roles