JobHabor

Sr. DevOps Engineer - AI and Site Reliability Engineering

Teradata

Location
California, USA
Workplace
Remote
Employment
Full Time
Salary
USD 121,900–182,800/yr
Apply on the employer’s site

Posted 2mo ago

The employer’s full description could not be read from their board. This is a summary of the posting — follow the apply link for the original.

Responsibilities

  • Design, implement, test, deploy, administer, and improve software solutions for system reliability and availability
  • Mitigate operational risks and track system health
  • Improve mean-time-to-discover and mean-time-to-respond for operational issues
  • Lead chaos engineering efforts in a production-alike environment
  • Leverage AI technologies (LLMs, ML, agentic systems) for operational efficiency and reliability improvements
  • Become a subject-matter expert in production deployment and upgrade of Teradata software
  • Work closely with product engineering and cloud operations personnel
  • Work with security and compliance teams to provide evidence for compliance obligations

Requirements

  • Bachelor’s degree or equivalent in computer science or a related field
  • 4+ years of industry experience
  • Experience with at least one major cloud service provider (AWS, Azure, and/or Google Cloud)
  • Experience building and deploying complex software solutions to significant operational problems
  • Proficiency with at least one modern programming language such as Python
  • Proficiency with a modern source control tool, preferably Git
  • Familiarity with machine learning libraries such as Tensorflow and Scikit-Learn
  • Experience building and deploying AI systems via cloud-based generative AI and agentic AI platforms
  • Experience with at least one modern defect tracking tool, preferably Jira
  • Experience with an infrastructure-as-code (IaC) cloud provisioning tool, preferably Terraform
  • Experience with a configuration management tool such as Ansible or Puppet
  • Experience with Grafana or an equivalent observability tool
  • Experience with a build/deployment automation tool such as Jenkins or Bamboo
  • Familiarity with both SQL and noSQL databases, and use cases for each
  • Experience administering Linux-based systems
  • 4+ years of experience in a devops or site reliability engineering role
  • An in-depth understanding of site reliability engineering principles
  • An understanding of enterprise software deployment and security/compliance principles
  • Proficiency with multi-layered technical troubleshooting and root-cause analysis
  • Ability to quickly and comprehensively decompose a problem, identifying dependencies and defining tasks
  • Ability to think creatively and holistically about solutions
  • Ability to work both independently and collaboratively in a fast-paced environment
  • Ability to adjust as priorities change
  • Ability to communicate concisely but effectively with colleagues, leaders, and stakeholders
  • Ability to tailor communications to the needs and understanding of a particular audience
  • Flexibility to work on a globally-distributed team managed from the United States

Preferred

  • Master’s degree or equivalent preferred
  • CSP developer or architect certifications preferred

Skills

  • Python
  • Git
  • Tensorflow
  • Scikit-Learn
  • AWS Bedrock
  • AWS Sagemaker
  • Azure AI Foundry
  • Google Vertex AI
  • Google AgentSpace
  • Jira
  • Terraform
  • Ansible
  • Puppet
  • Grafana
  • Jenkins
  • Bamboo
  • SQL
  • NoSQL
  • Linux
  • AI
  • Large Language Models
  • Machine Learning
  • Agentic Systems

Similar roles