JobHabor

Senior Site Reliability Engineer, Government

SentinelOne

Location
United States
Workplace
Remote
Employment
Full Time
Salary
USD 132,000–182,000/yr
Apply on the employer’s site

Posted 2mo ago

The employer’s full description could not be read from their board. This is a summary of the posting — follow the apply link for the original.

Responsibilities

  • Drive continuous software delivery, resolve incidents, run post mortems, and create automation strategies
  • Lead and execute incident management for production issues
  • Improve and optimize the observability strategy
  • Define, implement, and monitor SLOs, SLIs, and SLAs
  • Design, develop, and maintain software solutions for operational challenges
  • Own and coordinate all government environment releases
  • Understand product architecture and service dependencies
  • Partner cross-functionally with engineering, product, SecOps, compliance, and leadership teams
  • Ensure all infrastructure and deployments meet FedRAMP, government regulations, and industry standards

Requirements

  • 5+ years of experience in SRE, DevOps, or Infrastructure Engineering for SaaS products
  • 4+ years running operations at a large scale
  • 2+ years of production experience with a container orchestration system (Kubernetes preferred)
  • 2+ years of production experience with Continuous Delivery
  • Strong understanding of compliance frameworks relevant to government deployments (e.g., FedRAMP, DoD, NIST 800 53, NIST 800 137)
  • Multi cloud experience in AWS/GCP (expertise within AWS preferred)
  • Demonstrated experience with at least one main programming language (Python, Go, Ruby, etc.)
  • Proficiency in bash scripting to improve operational workflows
  • Familiarity with GitOps frameworks
  • Familiarity with IaC tooling (Terraform or Pulumi)
  • Familiarity with deployment strategies (blue green, rolling deploys, canary deploys)
  • Experience with industry standard observability stacks (Prometheus, Grafana, ELK, OpenTelemetry, etc.)
  • Experience with incident management processes
  • Proven background implementing and supporting FedRAMP, security, risk management, and compliance processes for software releases
  • Experience working directly with government agencies or in highly regulated industries
  • Familiarity with testing strategies and automation in large scale environments
  • U.S. Citizenship and a work location in the United States is required

Skills

  • Kubernetes
  • Python
  • Go
  • Ruby
  • Bash
  • GitOps
  • Terraform
  • Pulumi
  • Prometheus
  • Grafana
  • ELK
  • OpenTelemetry
  • AWS
  • GCP

Similar roles