JobHabor

Site Reliability Engineer

Datavant

Location
US
Workplace
Remote
Employment
Full Time
Salary
USD 100,000–130,000/yr
Apply on the employer’s site

Posted 1mo ago

The employer’s full description could not be read from their board. This is a summary of the posting — follow the apply link for the original.

Responsibilities

  • Lead the design and implementation of reliability improvements
  • Identify systemic inefficiencies and drive solutions
  • Implement customized solutions to complex operational problems
  • Review code, systems, and configuration for efficiency and best practices
  • Own SLO/SLI definitions and drive teams toward meeting targets
  • Lead incident response, facilitate postmortems, and ensure action items
  • Support the integration of newly acquired cloud environments
  • Contribute to hybrid-cloud and cross-cloud connectivity strategies
  • Help maintain and expand the modular network security edge
  • Drive standardization of infrastructure and operational practices
  • Partner with security and platform engineering teams
  • Develop tools and processes to improve team service delivery
  • Analyze service delivery data and team feedback
  • Participate in and help evolve on-call practices, runbooks, and alerting strategy
  • Address communication gaps and produce clear documentation
  • Teach and lead more junior engineers
  • Accept and promote sound engineering decisions
  • Collaborate cross-functionally to embed reliability thinking
  • Build and maintain Infrastructure as Code
  • Develop and improve CI/CD pipelines, automated testing, and deployment tooling
  • Implement policy-based automation
  • Extend observability coverage
  • Leverage AI tools and agents

Requirements

  • 5+ years of experience in site reliability engineering, DevOps, or platform/infrastructure engineering
  • Strong expertise in cloud infrastructure, including networking, security, compute, storage, and IAM
  • Experience supporting workload migrations and integrating new cloud environments
  • Hands-on experience with Infrastructure as Code and automation (Terraform and Ansible)
  • Strong security knowledge, including IAM, encryption, network security, and compliance frameworks (SOC2, HITRUST, NIST)
  • Demonstrated ability to solve complex operational problems independently and drive solutions end-to-end
  • Strong proficiency in at least one systems language (Python, Go, or similar) and comfort across multiple languages and configuration formats
  • Proven ability to conduct meaningful code and system reviews not just for correctness, but for efficiency and architectural soundness
  • Strong communication skills, including the ability to document technical decisions, address gaps in understanding, and build team alignment
  • Experience leveraging AI agents to accelerate daily workload
  • Competency with AWS EC2, Azure VMs, Kubernetes / containerized workloads
  • Competency with AWS (VPC, TGW, Peering, SG, ALB, NLB); Azure (VNets, NSG, AGW); VPN
  • Competency with Datadog, CloudWatch, Azure Log Analytics instrumentation, dashboards, alerting
  • Competency with Terraform, Ansible, GitHub Actions or equivalent CI/CD tooling
  • Hands-on experience in AWS and/or Azure; familiarity with GCP
  • Competency with AWS IAM, Azure RBAC, least-privilege access patterns
  • Competency with Linux (required), Windows (helpful); DNS / IPAM

Preferred

  • Experience with multi-cloud environments (AWS, Azure, GCP) and multi-account cloud governance (AWS Organizations, Azure Policy, SCPs)
  • Background in driving environment consolidation and integration bringing order to disparate hybrid and multi-cloud architectures
  • Experience defining and operating against SLOs/SLIs and error budgets
  • Familiarity with FinOps principles and cloud cost optimization
  • Track record of improving on-call health reducing alert fatigue, improving runbook quality, driving down MTTR
  • Experience with cloud-native identity management (AWS IAM Identity Center, Entra ID, RBAC)
  • Knowledge of multi-region architecture and disaster recovery strategies

Skills

  • Python
  • Go
  • Terraform
  • Ansible
  • AWS
  • Azure
  • GCP
  • Kubernetes
  • Docker
  • Datadog
  • CloudWatch
  • Azure Log Analytics
  • GitHub Actions
  • AWS EC2
  • Azure VMs
  • AWS VPC
  • AWS TGW
  • AWS Peering
  • AWS SG
  • AWS ALB
  • AWS NLB
  • Azure VNets
  • Azure NSG
  • Azure AGW
  • VPN
  • AWS IAM
  • Azure RBAC
  • Linux
  • Windows
  • DNS
  • IPAM
  • AWS Organizations
  • Azure Policy
  • SCP
  • AWS IAM Identity Center
  • Entra ID

Similar roles