Sr. DevOps Engineer - AI and Site Reliability Engineering
Teradata
- Location
- California, USA
- Workplace
- Remote
- Employment
- Full Time
- Salary
- USD 121,900–182,800/yr
Posted 2mo ago
The employer’s full description could not be read from their board. This is a summary of the posting — follow the apply link for the original.
Responsibilities
- Design, implement, test, deploy, administer, and improve software solutions for system reliability and availability
- Mitigate operational risks and track system health
- Improve mean-time-to-discover and mean-time-to-respond for operational issues
- Lead chaos engineering efforts in a production-alike environment
- Leverage AI technologies (LLMs, ML, agentic systems) for operational efficiency and reliability improvements
- Become a subject-matter expert in production deployment and upgrade of Teradata software
- Work closely with product engineering and cloud operations personnel
- Work with security and compliance teams to provide evidence for compliance obligations
Requirements
- Bachelor’s degree or equivalent in computer science or a related field
- 4+ years of industry experience
- Experience with at least one major cloud service provider (AWS, Azure, and/or Google Cloud)
- Experience building and deploying complex software solutions to significant operational problems
- Proficiency with at least one modern programming language such as Python
- Proficiency with a modern source control tool, preferably Git
- Familiarity with machine learning libraries such as Tensorflow and Scikit-Learn
- Experience building and deploying AI systems via cloud-based generative AI and agentic AI platforms
- Experience with at least one modern defect tracking tool, preferably Jira
- Experience with an infrastructure-as-code (IaC) cloud provisioning tool, preferably Terraform
- Experience with a configuration management tool such as Ansible or Puppet
- Experience with Grafana or an equivalent observability tool
- Experience with a build/deployment automation tool such as Jenkins or Bamboo
- Familiarity with both SQL and noSQL databases, and use cases for each
- Experience administering Linux-based systems
- 4+ years of experience in a devops or site reliability engineering role
- An in-depth understanding of site reliability engineering principles
- An understanding of enterprise software deployment and security/compliance principles
- Proficiency with multi-layered technical troubleshooting and root-cause analysis
- Ability to quickly and comprehensively decompose a problem, identifying dependencies and defining tasks
- Ability to think creatively and holistically about solutions
- Ability to work both independently and collaboratively in a fast-paced environment
- Ability to adjust as priorities change
- Ability to communicate concisely but effectively with colleagues, leaders, and stakeholders
- Ability to tailor communications to the needs and understanding of a particular audience
- Flexibility to work on a globally-distributed team managed from the United States
Preferred
- Master’s degree or equivalent preferred
- CSP developer or architect certifications preferred
Skills
- Python
- Git
- Tensorflow
- Scikit-Learn
- AWS Bedrock
- AWS Sagemaker
- Azure AI Foundry
- Google Vertex AI
- Google AgentSpace
- Jira
- Terraform
- Ansible
- Puppet
- Grafana
- Jenkins
- Bamboo
- SQL
- NoSQL
- Linux
- AI
- Large Language Models
- Machine Learning
- Agentic Systems
Similar roles
Staff AI Engineering - Enterprise Architecture
American Express · Phoenix, AZ, United States · USD 144,250–256,250/yr · today
Field Service Technician II
Toshiba America Business Solutions · East Syracuse, NY, United States · today
Field Service Technician I
Toshiba America Business Solutions · Rochester, NY, United States · today
Field Service Technician II
Toshiba America Business Solutions · Buffalo, NY, United States · today
Incentives Analyst
MicroStrategy · Tysons Corner, VIRGINIA, United States · USD 66,400–119,600/yr · today
Senior Azure Cloud Engineer
Vaxcyte · San Carlos, California, United States · today