JobHabor

Senior Site Reliability Engineer

PandaDoc
Location
Kyiv · Warsaw · Lisbon
Workplace
Remote
Employment
Full Time
Salary
Apply on the employer’s site

Posted 7mo ago

The employer’s full description could not be read from their board. This is a summary of the posting — follow the apply link for the original.

Responsibilities

  • Own and influence the incident management process end-to-end
  • Maintain and evolve on-prem observability stack
  • Keep production applications running smoothly by participating in the on-call rotation
  • Develop automations and tools to support platform reliability
  • Contribute to production services with performance and resiliency in mind
  • Collaborate with product engineers to foster SRE principles within the R&D organization
  • Be a mentor for the SRE team or product engineers

Requirements

  • Solid programming experience, namely Python (Django and AsyncIO) and/or Java (Spring Boot)
  • Experience in maintaining an observability tools suite (specifically, LGTM - Loki, Grafana, Tempo, Mimir)
  • Experience in development and maintenance of Python services in production
  • Strong experience with AWS and Kubernetes
  • Solid proficiency in working with relational databases (PostgreSQL)
  • Solid proficiency in working with messaging systems (e.g. RabbitMQ, NATS, Kafka)
  • An experienced on-call SRE engineer
  • Enjoy hands-on troubleshooting of distributed systems in production environments
  • You act like an owner and strive to do work you're proud of
  • You enjoy communication and knowledge sharing on all-things reliability
  • Proficiency in English, both written and spoken

Skills

  • Python
  • Django
  • AsyncIO
  • Java
  • Spring Boot
  • Loki
  • Grafana
  • Tempo
  • Mimir
  • AWS
  • Kubernetes
  • PostgreSQL
  • RabbitMQ
  • NATS
  • Kafka

More jobs at PandaDoc

All 7

Similar roles