JobHabor

Senior Site Reliability Engineer

Clickhouse
Location
US
Workplace
Remote
Employment
Full Time
Salary
USD 141,000–208,000/yr
Apply on the employer’s site

Posted 5mo ago

The employer’s full description could not be read from their board. This is a summary of the posting — follow the apply link for the original.

Responsibilities

  • Build and lead processes for reliability, availability, scalability, and performance of cloud infrastructure
  • Collaborate with teams to design and implement scalable, secure, highly available, and fault-tolerant distributed systems
  • Own incident management and response
  • Perform post-mortem analysis and run blameless postmortems
  • Continuously improve Cloud services
  • Leverage software engineering expertise to develop platforms and tools for operational and engineering efficiencies
  • Design and implement scalable, secure, and highly available systems for ClickHouse
  • Establish and manage service level objectives (SLOs) and service level agreements (SLAs)
  • Ensure infrastructure components have monitoring and alerting
  • Enhance and refine incident response processes and post-mortem analysis
  • Communicate with the support team to inform impacted customers
  • Continuously improve reliability and performance of ClickHouse services
  • Plan, enable, and drive Chaos initiatives
  • Manage on-call processes to respond to performance and reliability issues
  • Establish best practices for coordinating escalation to resolve issues and minimize downtime

Requirements

  • Bachelor’s or Master’s degree in Computer Science or a related field
  • At least 8 years of experience in Site Reliability Engineering or a related field
  • Hands-on experience with Go and/or Python
  • Strong knowledge of cloud computing platforms such as AWS, Azure, or Google Cloud Platform
  • Excellent understanding of distributed databases and SQL
  • Hands-on experience with container orchestration tools such as Kubernetes or Docker Swarm
  • Strong experience with automation and configuration management tools such as Ansible, Terraform, or Puppet
  • Strong problem solver with solid production debugging skills
  • Passionate about efficiency, availability, scalability, and data governance
  • Thrive in a fast-paced environment
  • See yourself as a partner with the business
  • High level of responsibility, ownership, and accountability
  • Excellent communication and interpersonal skills

Preferred

  • ClickHouse is a major plus

Skills

  • Go
  • Python
  • AWS
  • Azure
  • Google Cloud Platform
  • Kubernetes
  • Docker Swarm
  • Ansible
  • Terraform
  • Puppet
  • ClickHouse

More jobs at Clickhouse

All 62

Similar roles