JobHabor

Site Reliability Engineer

Ooma
Location
Dallas · Ashburn · San Jose
Workplace
Hybrid
Employment
Full Time
Salary
USD 120,000–170,000/yr
Apply on the employer’s site

Posted 17d ago

The employer’s full description could not be read from their board. This is a summary of the posting — follow the apply link for the original.

Responsibilities

  • Manage large data centers with hundreds of bare metal servers and VMs
  • Monitor and troubleshoot system performance, reliability, and availability
  • Manage server hardware lifecycle including firmware, BIOS, and RMA coordination
  • Administer virtualization platforms and storage appliances
  • Automate bare metal and VM provisioning and OS lifecycle at scale
  • Work hands-on with data center network infrastructure including VLANs, routing, and load balancers
  • Design and maintain scalable infrastructure using containers, Kubernetes, and microservices
  • Oversee configuration management with Ansible to eliminate configuration drift
  • Design and operate high-throughput Kafka clusters for event streaming
  • Implement name services including DNS, DHCP, NTP, and certificate management
  • Collaborate with development teams to influence system design choices
  • Participate in on-call rotations and conduct blameless post-mortems
  • Create comprehensive technical documentation and maintain knowledge bases

Requirements

  • 8+ years of experience as an SRE or related field
  • Bachelor's degree in Computer Science, Engineering, or related field
  • Extensive on-premises data center experience with hundreds of bare-metal servers and VMs
  • Hands-on server hardware experience including firmware and BIOS management
  • Deep knowledge of Linux operating systems including configuration and performance tuning
  • Strong understanding of Linux networking concepts and protocols
  • Experience with containers and orchestration, particularly Kubernetes on bare metal
  • Experience managing CI/CD pipelines with GitOps workflows, Argo CD, and Helm charts
  • Knowledge of observability tools such as Prometheus, ELK Stack, and Grafana
  • Experience with configuration management tools, particularly Ansible

Preferred

  • Advanced degree strongly preferred
  • Experience with multi-site or colocation deployments and cross-site disaster recovery
  • Infrastructure as code with Terraform, MAAS, Foreman, or similar
  • Scripting in Python, Bash, or Go
  • Experience operating Kafka or other high-throughput stateful distributed systems

Skills

  • Linux
  • Kubernetes
  • CI/CD
  • Ansible
  • Kafka
  • DNS
  • DHCP
  • NTP
  • VLAN
  • Routing
  • Load balancing
  • Prometheus
  • ELK Stack
  • Grafana
  • Argo CD
  • Helm
  • Terraform
  • MAAS
  • Foreman
  • Python
  • Bash
  • Go

More jobs at Ooma

Similar roles