Senior Site Reliability Engineer
Clickhouse- Location
- US
- Workplace
- Remote
- Employment
- Full Time
- Salary
- —
Posted 5mo ago
The employer’s full description could not be read from their board. This is a summary of the posting — follow the apply link for the original.
Responsibilities
- Build and lead processes for reliability, availability, scalability, and performance of cloud infrastructure
- Collaborate with teams to design and implement scalable, secure, highly available, and fault-tolerant distributed systems
- Own incident management and response
- Perform post-mortem analysis and run blameless postmortems
- Continuously improve Cloud services
- Develop software platforms and tools to optimize operational and engineering efficiencies
- Establish and manage service level objectives (SLOs) and service level agreements (SLAs)
- Ensure infrastructure components have monitoring and alerting
- Enhance and refine incident response processes
- Plan, enable, and drive Chaos initiatives
- Manage on-call processes and establish best practices for escalation
Requirements
- Bachelor's or Master's degree in Computer Science or related field
- At least 8 years of experience in Site Reliability Engineering or related field
- Hands-on experience with Go and/or Python
- Strong knowledge of cloud computing platforms such as AWS, Azure, or Google Cloud Platform
- Excellent understanding of distributed databases and SQL
- Hands-on experience with container orchestration tools such as Kubernetes or Docker Swarm
- Strong experience with automation and configuration management tools such as Ansible, Terraform, or Puppet
- Strong problem solver with solid production debugging skills
- Passionate about efficiency, availability, scalability, and data governance
- Thrive in a fast-paced environment
- High level of responsibility, ownership, and accountability
- Excellent communication and interpersonal skills
Preferred
- ClickHouse experience is a major plus
Skills
- Go
- Python
- AWS
- Azure
- Google Cloud Platform
- Kubernetes
- Docker Swarm
- Ansible
- Terraform
- Puppet
- ClickHouse
More jobs at Clickhouse
All 62Incident Response Security Engineer
Clickhouse · US · 23d ago
Incident Response Security Engineer
Clickhouse · Remote · USD 169,150–225,000/yr · 23d ago
AI Operations Engineer
Clickhouse · US · 1mo ago
Senior Cloud Software Engineer - Efficiency Engineering
Clickhouse · Global · USD 133,450–197,200/yr · 1mo ago
Senior Software Engineer - Python and Data Ecosystem
Clickhouse · USA · Canada · UK +10 · 1mo ago
Similar roles
Senior Business Engineer - Ads
Reddit · Remote · USD 180,200–252,300/yr · today
Machine Learning Engineer
Reddit · US · USD 185,800–303,400/yr · today
Firmware Engineer
Anduril Industries · Costa Mesa, California, United States · USD 166,000–220,000/yr · today
Embedded Firmware Engineer
Anduril Industries · Costa Mesa, California, United States · USD 166,000–220,000/yr · today
Firmware Engineer, Connected Warfare
Anduril Industries · Costa Mesa, California, United States · USD 132,000–198,000/yr · today
Hardware Integration & Test Engineer II
SEAKR Engineering · Centennial, CO, United States · today