Sr. Site Reliability Engineer
Donnelley Financial Solutions (DFIN)
- Location
- US
- Workplace
- Onsite
- Employment
- Full Time
- Salary
- —
Posted 1mo ago
Sr. Site Reliability Engineer-DFIN-US
DFIN
Sr. Site Reliability Engineer
Donnelley Financial Solutions (DFIN) is the leading provider of innovative, software and technology–enabled financial regulatory and compliance solutions. Supporting every stage of our clients’ business and investment lifecycle, DFIN provides regulatory filing and transactional deal solutions to public and private companies, mutual funds, and other regulated investment firms.
Apply Now View Summary
Map Learn More
Learn More
- Summary
- Overview
- Responsibilities
- Requirements
- About Us
DFIN
- SUMMARY
- Overview
- Responsibilities
- Requirements
- ABOUT US
- Apply Now
- Map
- Print to PDF
- Share via
Company
DFIN
Industry
Risk and Compliance Solutions
Level
Full Time
Job Family
Product Engineering
Location
US
Compensation
About Us
For more information
website
www.dfinsolutions.com
Overview
Position Summary
The Senior Site Reliability Engineer is responsible for ensuring our SaaS products are fast, stable and optimized for our customers. SRE’s at DFIN take on availability, performance, managing change, monitoring, response and are guardians of non-functional requirements.
Position Details
You either have an SaaS infrastructure background with a programmatic, automated mindset or are someone that comes with a software engineering background with SaaS infrastructure experience. The SRE goal is to build automated systems that reduce or eliminate manual work to keep our products up and running and performing optimally. We are looking for someone who thrives on collaboration within the team and across other groups and can operate independently to deliver solutions.
Responsibilities
Performance
Dive deep into technology and stay on the forefront of the latest tools, technologies, and strategies; help evaluate, prototype, and integrate them into work processes. Build strong relationships with SRE team members and software engineering teams to hold each other accountable for quality expectations. Learn continuously and apply lessons learned. Evangelize best practices, eliminate bottlenecks, and improve process.
Quality Assurance
Champion and implement a culture of SRE to maintain a high-quality platform infrastructure in DFIN SaaS products. Leverage AI tools to enhance system reliability, including intelligent observability, incident prediction and automated remediation across cloud infrastructure. Perform with broad independence and deliver on project milestones and tasks on schedule while communicating progress regularly.
Technology
Evaluate and implement emerging AI powered operations and observability solutions to proactively improve system performance, reliability and scalability. Champion and implement application and infrastructure monitoring and alerting to prevent client impacting issues by ensuring system availability, performance and scalability to maintain SLOs and SLAs. Automate everything, including system operational runbooks.
Support
Participate in on-call duties 365/24/7 and lead the triage and RCA of production incidents. Optimize application performance at scale. Define and support continuous integration and deployment pipelines (CI/CD) aligned to branching and quality assurance strategies.
Requirements
Education & Experience
BS in Computer Science or equivalent work experience. 5+ years' experience designing, building, securing, monitoring and maintaining cloud infrastructure in Azure or AWS. Experience applying AI capabilities within CloudOps operations. Relevant certifications or training in AI, Cloud AI services or AIOps platforms are a plus. 5+ years' experience writing scripts in PowerShell or Python/Bash to automate system operations as runbooks for Windows or Linux environments.
Proficiency
5+ years' experience writing software in any modern software language such as C# .NET, Java. Experience planning, coordinating, developing and executing all stages of post deployment verification test scripts. 5+ years' experience implementing production performance, availability, and scalability monitoring and alerting using a tool such as New Relic, Dynatrace, DataDog or AppDynamics. 5+ years' experience supporting public client facing revenue generating systems.
Production
5+ years' experience creating automated deployments with tools such as Harness, Azure DevOps, Ansible or Jenkins to manage Infrastructure as Code and software build and deployment in a continuous integration (CI) / continuous delivery (CD) environment. Experience securing Windows or Linux systems in 24x7 production environment.
Skills
Strong DevOps focus and experience building and deploying Infrastructure as Code with Terraform or similar technology. Experiencing monitoring and preventing issues with databases and database queries (SQL, Cosmos) using tools like Solarwinds Database Performance Analyzer, Idera SQL Diagnostic Manager, or Redgate SQL Monitor. Experience with containerization and managing Kubernetes clusters (AKS or EKS). Experience with common cloud networking, firewall and load balancing configuration.
About Us
For more information
website
www.dfinsolutions.com
Apply Now
keyboard_arrow_down
Print to pdf
Sr. Site Reliability Engineer
Donnelley Financial Solutions (DFIN) is the leading provider of innovative, software and technology–enabled financial regulatory and compliance solutions. Supporting every stage of our clients’ business and investment lifecycle, DFIN provides regulatory filing and transactional deal solutions to public and private companies, mutual funds, and other regulated investment firms.
View Summary
Company
DFIN
Location
US
Level
Full Time
Compensation
Job Family
Product Engineering
Industry
Risk and Compliance Solutions
Overview
We are looking for technical team members at all levels who want to push themselves to deliver best in market SaaS solutions. We offer a challenging environment where you will have to grow, adapt and use your skills consistently. Our customers rely on us in the moments that matter. Engineering delivers on that promise. The Senior Site Reliability Engineer is responsible for ensuring our SaaS products are fast, stable and optimized for our customers. SRE’s at DFIN take on availability, performance, managing change, monitoring, response and are guardians of non-functional requirements. You either have an SaaS infrastructure background with a programmatic, automated mindset or are someone that comes with a software engineering background with SaaS infrastructure experience. The SRE goal is to build automated systems that reduce or eliminate manual work to keep our products up and running and performing optimally. We are looking for someone who thrives on collaboration within the team and across other groups and can operate independently to deliver solutions.
read more We are looking for technical team members at all levels who want to push themselves to deliver best in market SaaS solutions. We offer a challenging environment where you will have to grow, adapt and use your skills consistently. Our customers rely on us in the moments that matter. Engineering delivers on that promise.
The Senior Site Reliability Engineer is responsible for ensuring our SaaS products are fast, stable and optimized for our customers. SRE’s at DFIN take on availability, performance, managing change, monitoring, response and are guardians of non-functional requirements.
You either have an SaaS infrastructure background with a programmatic, automated mindset or are someone that comes with a software engineering background with SaaS infrastructure experience. The SRE goal is to build automated systems that reduce or eliminate manual work to keep our products up and running and performing optimally. We are looking for someone who thrives on collaboration within the team and across other groups and can operate independently to deliver solutions.
Responsibilities
- Champion and implement a culture of SRE to maintain a high-quality platform infrastructure in DFIN SaaS products.
- Leverage AI tools to enhance system reliability, including intelligent observability, incident prediction and automated remediation across cloud infrastructure.
- Evaluate and implement emerging AI powered operations and observability solutions to proactively improve system performance, reliability and scalability.
- Champion and implement application and infrastructure monitoring and alerting to prevent client impacting issues by ensuring system availability, performance and scalability to maintain SLOs and SLAs.
- Optimize application performance at scale.
- Automate everything, including system operational runbooks.
- Define and support continuous integration and deployment pipelines (CI/CD) aligned to branching and quality assurance strategies.
- Dive deep into technology and stay on the forefront of the latest tools, technologies, and strategies; help evaluate, prototype, and integrate them into work processes.
- Perform with broad independence and deliver on project milestones and tasks on schedule while communicating progress regularly.
- Build strong relationships with SRE team members and software engineering teams to hold each other accountable for quality expectations.
- Learn continuously and apply lessons learned.
- Evangelize best practices, eliminate bottlenecks, and improve process.
- Participate in on-call duties 365/24/7 and lead the triage and RCA of production incidents.
read more
- Champion and implement a culture of SRE to maintain a high-quality platform infrastructure in DFIN SaaS products.
- Leverage AI tools to enhance system reliability, including intelligent observability, incident prediction and automated remediation across cloud infrastructure.
- Evaluate and implement emerging AI powered operations and observability solutions to proactively improve system performance, reliability and scalability.
- Champion and implement application and infrastructure monitoring and alerting to prevent client impacting issues by ensuring system availability, performance and scalability to maintain SLOs and SLAs.
- Optimize application performance at scale.
- Automate everything, including system operational runbooks.
- Define and support continuous integration and deployment pipelines (CI/CD) aligned to branching and quality assurance strategies.
- Dive deep into technology and stay on the forefront of the latest tools, technologies, and strategies; help evaluate, prototype, and integrate them into work processes.
- Perform with broad independence and deliver on project milestones and tasks on schedule while communicating progress regularly.
- Build strong relationships with SRE team members and software engineering teams to hold each other accountable for quality expectations.
- Learn continuously and apply lessons learned.
- Evangelize best practices, eliminate bottlenecks, and improve process.
- Participate in on-call duties 365/24/7 and lead the triage and RCA of production incidents.
Requirements
- 5+ years' experience designing, building, securing, monitoring and maintaining cloud infrastructure in Azure or AWS.
- Experience applying AI capabilities within CloudOps operations.
- Relevant certifications or training in AI, Cloud AI services or AIOps platforms are a plus.
- 5+ years' experience writing software in any modern software language such as C# .NET, Java.
- 5+ years' experience creating automated deployments with tools such as Harness, Azure DevOps, Ansible or Jenkins to manage Infrastructure as Code and software build and deployment in a continuous integration (CI) / continuous delivery (CD) environment.
- 5+ years' experience implementing production performance, availability, and scalability monitoring and alerting using a tool such as New Relic, Dynatrace, DataDog or AppDynamics.
- 5+ years' experience writing scripts in PowerShell or Python/Bash to automate system operations as runbooks for Windows or Linux environments.
- 5+ years' experience supporting public client facing revenue generating systems.
- Strong DevOps focus and experience building and deploying Infrastructure as Code with Terraform or similar technology.
- Experiencing monitoring and preventing issues with databases and database queries (SQL, Cosmos) using tools like Solarwinds Database Performance Analyzer, Idera SQL Diagnostic Manager, or Redgate SQL Monitor.
- Experience planning, coordinating, developing and executing all stages of post deployment verification test scripts.
- Experience securing Windows or Linux systems in 24x7 production environment.
- Experience with containerization and managing Kubernetes clusters (AKS or EKS).
- Experience with common cloud networking, firewall and load balancing configuration.
- BS in Computer Science or equivalent work experience.
read more
- 5+ years' experience designing, building, securing, monitoring and maintaining cloud infrastructure in Azure or AWS.
- Experience applying AI capabilities within CloudOps operations.
- Relevant certifications or training in AI, Cloud AI services or AIOps platforms are a plus.
- 5+ years' experience writing software in any modern software language such as C# .NET, Java.
- 5+ years' experience creating automated deployments with tools such as Harness, Azure DevOps, Ansible or Jenkins to manage Infrastructure as Code and software build and deployment in a continuous integration (CI) / continuous delivery (CD) environment.
- 5+ years' experience implementing production performance, availability, and scalability monitoring and alerting using a tool such as New Relic, Dynatrace, DataDog or AppDynamics.
- 5+ years' experience writing scripts in PowerShell or Python/Bash to automate system operations as runbooks for Windows or Linux environments.
- 5+ years' experience supporting public client facing revenue generating systems.
- Strong DevOps focus and experience building and deploying Infrastructure as Code with Terraform or similar technology.
- Experiencing monitoring and preventing issues with databases and database queries (SQL, Cosmos) using tools like Solarwinds Database Performance Analyzer, Idera SQL Diagnostic Manager, or Redgate SQL Monitor.
- Experience planning, coordinating, developing and executing all stages of post deployment verification test scripts.
- Experience securing Windows or Linux systems in 24x7 production environment.
- Experience with containerization and managing Kubernetes clusters (AKS or EKS).
- Experience with common cloud networking, firewall and load balancing configuration.
- BS in Computer Science or equivalent work experience.
Apply Now
Close
Watch
Name* Name should have minimum 2 and maximum 60 characters Name is a required field Name should have minimum 2 and maximum 60 characters
Email Address* Email is a required field Please enter a valid email address
Phone Number* Phone number is a required field For non-US/Canadian phone numbers please use + country code
Please enter a valid phone number
LinkedIn Profile Please enter a valid linkedin address
Upload Resume Please upload file in PDF, DOC, DOCX, TXT or RTF format up to 10 MB Please select file less than 10MB. Please upload file in PDF, DOC, DOCX, TXT or RTF format
Message Message should have minimum 10 and maximum 5000 characters Message should have minimum 10 and maximum 5000 characters
Apply
You have successfully applied for this Vizi!
Oops...something went wrong!
Skills
- AWS
- Azure
- GCP
- Python
- Bash
- Docker
- Kubernetes
- Prometheus
- Grafana
- ELK stack
- Terraform
- Ansible
- CI/CD
Similar roles
Senior Business Engineer - Ads
Reddit · Remote · USD 180,200–252,300/yr · today
Machine Learning Engineer
Reddit · US · USD 185,800–303,400/yr · today
Firmware Engineer
Anduril Industries · Costa Mesa, California, United States · USD 166,000–220,000/yr · today
Embedded Firmware Engineer
Anduril Industries · Costa Mesa, California, United States · USD 166,000–220,000/yr · today
Firmware Engineer, Connected Warfare
Anduril Industries · Costa Mesa, California, United States · USD 132,000–198,000/yr · today
Hardware Integration & Test Engineer II
SEAKR Engineering · Centennial, CO, United States · today