Digital System Support Engineer

We are currently recruiting for a Digital System Support Engineer based in the UK. These engineers will work 10-hour shifts on a 4-day working week, providing coverage seven days a week (Monday through Sunday). Together, they will ensure full weekly coverage. Shift schedules can either be fixed or rotated, depending on the engineers' preferences.

These roles do not require night shifts. The engineers will not be required to participate in an on-call rotation.

Shift Pattern

  • 4-day working week with 10-hour shifts

  • Coverage across Monday to Sunday (shared between two engineers)

  • Shifts typically begin around 7:00 AM — no night shifts required

  • Fixed or rotating schedule based on preference

  • No on-call rotation required

Benefits

  • 26.5 days annual leave plus your birthday off

  • Salary sacrifice company pension scheme

  • Personal life insurance (3x your salary) and income protection

  • Health insurance with options to add your family

  • Dedicated professional development training budget

  • Enhanced company sick pay

  • Enhanced family leave policy

  • 2 paid volunteer days per year

  • Electric Car Scheme (salary sacrifice)

As a Digital System Support Engineer, you will:

  • Provide real-time monitoring of all Caesars Digital production environments including online sportsbook, iCasino, retail platforms, and supporting infrastructure across multiple monitoring platforms (Zabbix, Splunk, DynaTrace, New Relic).

  • Serve as the first line of defense for service assurance by performing first-touch validation of all incoming alerts, classifying severity based on impact and urgency, and determining whether alerts are actionable or false positives.

  • Execute documented runbook procedures for known alerts, resolving issues where possible, and routing to the correct resolver groups when escalation is required.

  • Proactively identify anomalies and performance degradation before automated alerts fire by recognising unusual patterns on dashboards and correlating multiple signals that may indicate a broader systemic issue.

  • Initiate major incident communication protocols for Critical incidents including creating bridge channels, paging the Major Incident Manager on-call, and posting initial stakeholder notifications.

  • Monitor business service dashboards including customer journey health.

  • Escalate immediately when business metrics deviate from expected baselines, including revenue-impacting degradation, even in the absence of a technical alert.

  • Execute shift handoff protocols by thoroughly documenting active issues, pending escalations, upcoming change windows, and items requiring continued attention for the incoming shift.

  • Maintain situational awareness of all planned changes, maintenance windows, and deployment activities that may affect service availability.

  • Validate synthetic monitoring results and flag service degradation, ensuring customer-facing services remain reachable and performant.

  • Report noisy and non-actionable alert patterns to the Digital System Support Lead for threshold tuning and alert optimization.

  • Update Jira tickets with accurate triage notes, timeline entries, and status changes to maintain proper incident documentation.

  • Identify and flag recurring alert patterns and "near-miss" incidents that were caught before customer impact, contributing to continuous improvement of detection capabilities.

  • Provide business impact context in incident communications by correlating customer journey and revenue dashboard data with infrastructure alerts.

  • Support enhanced monitoring postures during peak events (Super Bowl, March Madness, NFL season openers, major sporting events, state launches) as coordinated by the Major Incident Manager.

Education and Experience

  • 2-4 years of experience in application support, IT operations, Network Operations Center (NOC), or a similar operational environment.

  • Hands-on experience with enterprise monitoring tools such as Zabbix, Splunk, DynaTrace, New Relic, or similar observability platforms.

  • Strong understanding of incident management processes including alert triage, severity classification, and escalation procedures.

  • Excellent problem-solving and analytical skills with the ability to correlate multiple data points to identify root issues.

  • Strong written and verbal communication skills, particularly in high-pressure situations requiring clear and concise stakeholder updates.

  • Experience working in a fast-paced, high-pressure environment while handling multiple simultaneous tasks and priorities.

  • Understanding of cloud-based technologies, microservices architectures, and distributed systems.

  • Familiarity with ITIL processes, particularly Incident Management and Event Management.

  • Previous experience in gaming, sports betting, or digital entertainment industries is preferred.

  • Understanding of technical infrastructure including networking, APIs, databases, and data workflows.

  • Associate or Bachelor's degree in Computer Science, Information Technology, or a related field, or equivalent work experience.

See also

要針對這個職缺調整履歷嗎?

目前無法檢查您與這個職缺的符合程度;請先將履歷加入個人檔案,下次即可查看。

A new version of freehire is available