JobHabor

AI Operations and Monitoring Engineer

Choctaw Nation of Oklahoma
Location
Durant, OK, United States
Workplace
Hybrid
Employment
Full Time
Salary
Apply on the employer’s site

Posted 15d ago

Monday-Friday 8:00AM-4:30PM| Hybrid Position| Weekly Earned Wage Access is an option for this position.

Job Purpose or Goals

The AI Operations and Monitoring Engineer is responsible for ensuring the reliability, performance, documentation, and compliance of production AI/ML systems, keeping them stable and functioning as intended. They play a key role in maintaining system uptime and driving effective incident response to minimize outages and protect overall service quality.

Tasks

  • Monitor AI/ML system health, drift, and anomalies by continuously tracking model performance, identifying unexpected behavior, and ensuring systems operate within expected parameters.
  • Build dashboards, alerts, and observability tooling to give teams real-time insight into system performance and enable rapid detection of issues before they impact operations.
  • Implement CI/CD pipelines for models to streamline deployments, maintain version control, and ensure that model updates move reliably from development to production.
  • Maintain training and deployment environments by keeping infrastructure stable, updated, and optimized so models can be trained and deployed efficiently.
  • Develop automation for scaling, retraining and reducing manual workloads and ensuring models automatically adapt to new data or demand.
  • Respond to incidents and outages to restore services quickly, conduct root cause analysis, and prevent future disruptions to AI/ML workloads.
  • Collaborate on security, compliance, and governance to ensure AI systems follow organizational standards and regulatory requirements.
  • Write runbooks and maintain audit logs to document procedures, support operational consistency, and provide traceability for all system actions.
  • Performs other duties as may be assigned.

Job Requirements

Bachelor's degree in computer science or related field, or 4 years relevant professional experience

3 years' experience in DevOps, SRE, or MLOps Proficiency with tools like Kubernetes, Docker, and cloud services

Experience with observability tools

Familiarity with ML deployment workflows

  • Monitor AI/ML system health, drift, and anomalies by continuously tracking model performance, identifying unexpected behavior, and ensuring systems operate within expected parameters.
  • Build dashboards, alerts, and observability tooling to give teams real-time insight into system performance and enable rapid detection of issues before they impact operations.
  • Implement CI/CD pipelines for models to streamline deployments, maintain version control, and ensure that model updates move reliably from development to production.
  • Maintain training and deployment environments by keeping infrastructure stable, updated, and optimized so models can be trained and deployed efficiently.
  • Develop automation for scaling, retraining and reducing manual workloads and ensuring models automatically adapt to new data or demand.
  • Respond to incidents and outages to restore services quickly, conduct root cause analysis, and prevent future disruptions to AI/ML workloads.
  • Collaborate on security, compliance, and governance to ensure AI systems follow organizational standards and regulatory requirements.
  • Write runbooks and maintain audit logs to document procedures, support operational consistency, and provide traceability for all system actions.
  • Performs other duties as may be assigned.

Bachelor's degree in computer science or related field, or 4 years relevant professional experience

3 years' experience in DevOps, SRE, or MLOps Proficiency with tools like Kubernetes, Docker, and cloud services

Experience with observability tools

Familiarity with ML deployment workflows

Skills

  • MLOps
  • Kubernetes
  • Docker
  • Machine Learning

More jobs at Choctaw Nation of Oklahoma

All 10

Similar roles