Incident Manager
Caesars Entertainment- Location
- Leeds, West Yorkshire, United Kingdom
- Workplace
- Hybrid
- Employment
- Full Time
- Salary
- GBP 51,000–70,000/yr
Posted 16d ago
About the Role
We are recruiting an Incident Manager based in the UK. The Incident Manager will work 8-hour shifts across a five-day week and provide on-call assistance as part of a four-person rota that delivers 24/7 escalation coverage. Shift schedules may be fixed or rotated, depending on engineers' preferences.
These roles do not require night shifts.
Benefits
- 33 days annual leave plus your birthday day
- Salary sacrifice company pension scheme
- Personal life insurance (3x your salary) and income protection
- Health insurance with options to add your family
- Dedicated professional development training budget
- Enhanced company sick pay
- Enhanced family leave policy
- 2 paid volunteer days per year
- Electric Car Scheme (salary sacrifice)
Salary Range
£51,000 - £70,000
As a Digital Incident Manager, you will
- Assume the Incident Commander role for all P1 and P2 incidents: own the bridge call, drive resolution, and coordinate cross-functional resolver groups.
- Establish roles on incident bridges (scribe, technical lead, communications lead) and enforce time-boxed troubleshooting with 30-minute checkpoints to avoid stagnation.
- Make escalation decisions: page on-call engineers, engage leadership or vendors, notify compliance, or execute rollback/failover/service-degradation procedures.
- Send initial stakeholder notifications within the SLA and maintain a consistent cadence: technical details for engineering; business impact for leadership.
- Coordinate with Compliance and Regulatory teams for impact notifications and initiate customer-facing communications (app banners, status page updates) when required.
- Enforce monitoring coverage requirements for all Caesars Digital production services, ensure no service goes to production without adequate monitoring and alerting in place.
- Continuously tune alert thresholds based on feedback from Digital System Support Engineers, reducing false positives and alert fatigue while driving the alert signal-to-noise ratio above 80% actionable.
- Implement alert deduplication, correlation, and suppression rules to ensure Engineers receive clean, actionable signals rather than noise that degrades response effectiveness.
- Define and maintain alert severity standards that clearly distinguish P1 vs. P2 vs. P3 vs. informational alerts, ensuring consistent classification across all services.
- Review "missed detection" findings from the Major Incident Manager's Post-Incident Reviews and build new monitoring coverage to prevent recurrence of undetected issues.
- Ensure every alert in the ecosystem links to a documented runbook with clear response procedures that Engineers can execute independently.
- Build and own end-to-end customer journey monitoring covering the critical user flows: Registration → Deposit → Bet Placement → Bet Settlement → Withdrawal, with defined thresholds for success rates and drop-off alerts.
- Design and implement real-time revenue monitoring dashboards tracking deposit/withdrawal volumes, payment gateway health, and transaction success rates with anomaly detection against expected baselines.
- Build revenue impact calculation models for use during major incidents, enabling the team to quantify business impact in dollar terms.
- Define and maintain business KPI dashboards monitoring operational metrics including handle, active users, concurrent sessions, bet volume per minute, and Caesars Rewards pipeline health.
- Create executive-visible business health dashboards that provide real-time situational awareness during high-revenue events and peak traffic periods.
- Conduct monthly monitoring audits to assess coverage completeness, alert quality, and identify stale or orphaned alerts for decommissioned services.
- Build synthetic monitoring for critical customer journeys to validate service availability and performance from the customer's perspective.
- Collaborate with the Problem Manager to support event readiness by building enhanced dashboards, lowering detection thresholds, and adding event-specific synthetic monitors per the readiness plan.
- Report on noisy alert sources monthly and drive engineering teams to fix the root causes generating non-actionable alerts.
- Partner with product and engineering teams to agree on monitoring thresholds and ensure observability is built into the development lifecycle.
EDUCATION AND EXPERIENCE
- 5+ years of experience in IT operations, site reliability engineering, observability engineering, or a senior technical monitoring role.
- Deep hands-on expertise with enterprise monitoring and observability platforms such as Zabbix, Splunk, DynaTrace, New Relic, Datadog, Grafana, or similar tools.
- Proven experience designing and implementing monitoring strategies for large-scale, distributed, customer-facing digital platforms.
- Strong understanding of alert engineering principles including threshold tuning, correlation rules, suppression logic, and noise reduction techniques.
- Experience building business-level monitoring including revenue dashboards, customer journey tracking, and transaction anomaly detection.
- Ability to translate technical metrics into business impact language for executive stakeholders.
- Strong analytical skills with experience in defining KPIs, setting baselines, and measuring Mean Time to Detect (MTTD) improvements.
- Experience working with engineering teams to embed observability into CI/CD pipelines and service delivery processes.
- Understanding of cloud-native architectures, microservices, containerization, and modern infrastructure patterns.
- Previous experience in gaming, sports betting, or high-transaction digital commerce environments is strongly preferred.
- Familiarity with ITIL processes, particularly Event Management, and experience with on-call escalation models.
- Bachelor's degree in Computer Science, Information Technology, Engineering, or a related field, or equivalent professional experience.
Skills
- Zabbix
- Splunk
- Dynatrace
- New Relic
- Datadog
- Grafana
More jobs at Caesars Entertainment
All 31Sr. Analytics Manager - Enterprise Data Solutions (EDS) - Corporate (Las Vegas)
Caesars Entertainment · Las Vegas, NV, United States · 3d ago
Analyst I POS (Point of Sale) - Las Vegas
Caesars Entertainment · Las Vegas, NV, United States · 3d ago
Senior IT Architect
Caesars Entertainment · Las Vegas, NV, United States · 3d ago
Senior Frontend Engineer - Acquisition
Caesars Entertainment · Jersey City, NJ, United States · Las Vegas, NV, United States · USD 175,000–180,000/yr · 4d ago
IT Property Support ENGINEER I
Caesars Entertainment · Robinsonville, MS, United States · 7d ago
Similar roles
Senior Network Engineer
Jane Street · London, England, United Kingdom · today
Technical Support Engineer II - W&B EMEA
CoreWeave · London, England · today
Platform and Services Engineer - IAM (we have office locations in Cambridge, Leeds and London)
Genomics England · London, England, United Kingdom · today
Technical Support (Telecoms) - 2nd Line
Radius · Crewe, England, United Kingdom · today
AI / Machine Learning Engineer Leader
Prolific Academic · UK · USD 200/hr · today
Technical Program Manager - International
MS Career - Contingent Worker · United Kingdom · today