JobHabor

Department Manager of SRE/AIOps Platforms

ConEd
Location
New York, NY, United States
Workplace
Employment
Full Time
Salary
USD 165,000
Apply on the employer’s site

Posted today

As the Manager of Site Reliability Engineering (SRE) and AI IT Operations (AIOps) Platforms at Con Edison, you will lead the strategy, implementation, and continuous evolution of the companys SRE and AIOps platforms, practices, and transformation initiatives. You will be responsible for developing automation across IT Operations by leveraging Agentic AI, observability, and industry-leading reliability engineering practices, with the long-term goal of transforming the organization to an SRE operating model. This includes modernizing and automating IT Operations, the Service Desk, NOC, Infrastructure Operations, Incident and Major Incident Response, while continuously improving reliability, reducing outages and MTTR, and enhancing the employee experience. You will build, mentor, and develop high-performing teams, foster a culture of operational excellence and continuous learning, and manage budgets, vendor partnerships, and technology investments to maximize business value. Ultimately, you will empower Con Edison employees to serve customers and the people of New York City through a highly reliable, secure, performant, and intuitive technology environment.

This position does not provide employment pursuant to the terms of a STEM OPT Training Plan.

Core Responsibilities

  • Lead the implementation and operationalization of Con Edison's Site Reliability Engineering, Observability, Digital Twin, Automation, and AIOps capabilities.
  • Drive a culture of innovation and creativity, serving as an incubator for emerging technologies and top engineering talent that elevates automation and AI capabilities across the enterprise.
  • Manage multiple capital and O&M budgets, appropriation requests, and purchase order lifecycles, ensuring strict alignment with regulatory rulings and financial governance.
  • Lead cross-functional teams to govern tool selection and roadmap execution, establishing unified SRE and observability platforms with advanced capabilities such as Digital Twin, end-to-end correlation, and predictive analytics, to drive proactive monitoring for the NOC, Infra Ops, and Application Support.
  • Create the foundations and scale the AI Ops practice and Agentic AI platform to redefine IT Operations. Manage and scale modern AI-powered Contact Center platforms in the cloud.
  • Recruit, train, and mentor a team of Site Reliability Engineers and AI Developers, with a goal to transform IT Operations into an SRE operating model.
  • Drive process engineering and optimization to align IT operations with AIOps capabilities and industry frameworks (e.g., ITIL) for scalable, modern service delivery.
  • Manage MSP vendor(s) involved in SRE Platforms implementation, establishing a solid KPI framework and a customer-centric approach targeting high user satisfaction scores.
  • Engage externally with leading technology partners and participate regularly in conferences to proactively identify new trends and adjust the portfolio technology roadmap.
  • Lead, coach, and develop a high-performing team, fostering an inclusive environment of continuous growth and rallying staff around a forward-looking vision powered by AI and modern technology.
  • Ensure compliance with regulations and laws related to data privacy, critical infrastructure and cyber security.
  • Plan, coordinate, and execute projects and initiatives supporting the broader strategic goals of the business.

Required Education/Experience

  • Bachelor's Degree in Computer Science, Information Technology, Engineering or a related field and 12 years of related work experience or
  • Master's Degree in Computer Science, Information Technology, Engineering or a related field and 10 years of related work experience.

Preferred Education/Experience

  • Bachelor's Degree and 10 years of related work experience working in customer communications, back office program management, billing and case management related field work. Experience working in the Clean Energy Marketplace.

Relevant Work Experience

  • Extensive experience in IT Automation Implementation (former developer profile), with a proven track record of leadership, required.
  • Strong knowledge of IT network, infrastructure, cloud and digital workplace, required.
  • Demonstrated success in mentoring, upskilling, and incubating engineering talent, required.
  • Experience with AI and AI Ops platforms required, preferably in the cloud, preferred.
  • Experience with SRE model and implementation with IT teams and leading enterprise transformation, preferred.
  • Experience leading 24x7 IT operations, preferred.
  • Experience managing Capital and O&M budgets, purchase orders, appropriation requests, and financial governance, preferred.
  • Relevant experience with cloud, SRE and AI certifications, preferred.

Skills and Abilities

  • Demonstrated analytical skills
  • Strong verbal communication and listening skills

Licenses and Certifications

  • Driver's License Required

Physical Demands

  • Sit or stand to answer a phone for the duration of the workday
  • Ability to stoop, bend, reach, and kneel throughout the workday
  • Ability to read small print and symbols

Additional Physical Demands

  • The selected candidate will be assigned a System Emergency Assignment (i.e., an emergency response role) and will be expected to work non-business hours during emergencies, which may include nights, weekends, and holidays.
  • Travel as necessary

Skills

  • Workday

More jobs at ConEd

All 19

Similar roles