JobHabor

Site Reliability Engineer, Core Streaming

Yelp

Location
United States
Workplace
Remote
Employment
Full Time
Salary
USD 141,000–216,000/yr
Apply on the employer’s site

Posted 2mo ago

The employer’s full description could not be read from their board. This is a summary of the posting — follow the apply link for the original.

Responsibilities

  • Own the reliability, scalability, and operational health of Kafka clusters
  • Build and maintain automation for cluster operations, upgrades, capacity scaling, and incident recovery
  • Partner with engineering teams to enable new streaming use cases, advise on best practices, and ensure data pipeline reliability
  • Troubleshoot complex issues affecting data flow, performance, or stability, and lead root cause analyses
  • Execute Kafka version upgrades and platform migrations with minimal disruption to critical services
  • Participate in on-call rotations

Requirements

  • Solid SRE or infrastructure engineering foundation
  • Experience with infrastructure-as-code (Terraform)
  • Experience with configuration management (Puppet, Ansible, or equivalent)
  • Experience with cloud platforms (AWS preferred)
  • Experience with Linux operations
  • Production level experience with Kafka or similar technologies at scale including cluster upgrades, migrations, and capacity planning
  • Programming proficiency in Python, Java, or similar for tooling and automation
  • Strong debugging and systems-thinking skills, comfortable tracing data flow issues end-to-end across distributed systems

Preferred

  • Experience with Apache Flink or other stream processing frameworks
  • Familiarity with Kafka Client APIs (Producer, Consumer, Streams)
  • Experience building internal self-service tooling or developer platforms
  • Experience with incident response and management

Skills

  • Kafka
  • Flink
  • Terraform
  • Puppet
  • Ansible
  • AWS
  • Linux
  • Python
  • Java

Similar roles