Senior Site Reliability Engineer, Observability

This engineering-first role focuses on hands-on observability and reliability engineering across predominantly Windows-based Azure and AWS infrastructure supporting enterprise treasury customers. The role includes monitoring and alerting in New Relic, SLO/SLI and error-budget management, Terraform infrastructure as code, Azure DevOps pipelines, Incident.IO administration, incident response, and coaching engineering teams.

Responsibilities

  • Design and implement monitoring, alerting, and dashboards in New Relic across Azure and AWS
  • Write NRQL queries for troubleshooting, analysis, and reporting
  • Define and implement SLOs, SLIs, and error budgets
  • Lead alert noise reduction and signal quality engineering
  • Optimize observability costs through log ingestion management and pipeline rules
  • Develop and maintain Terraform infrastructure as code for monitoring and observability resources
  • Establish IaC governance standards
  • Author and troubleshoot Azure DevOps pipelines
  • Administer Incident.IO, including alert routing and notification workflows
  • Build incident management foundations, postmortem processes, on-call rotations, escalation policies, and response playbooks
  • Track MTTR, MTTD, and incident frequency
  • Respond to and debrief on production incidents
  • Enable engineering teams through workshops, consultation, documentation, and training

Requirements

  • 7+ years in Site Reliability Engineering, DevOps, or Platform Engineering focused on observability and production operations
  • Hands-on engineering experience combined with coaching and mentoring ability
  • Experience in Agile/Scrum environments
  • Expert New Relic and NRQL experience
  • Knowledge of structured logging, metrics collection, distributed tracing, dashboards, and alerts
  • Expertise with SLOs, SLIs, and error budgets
  • Experience with Incident.IO, PagerDuty, OpsGenie, or similar platforms
  • Experience designing incident workflows, on-call rotations, escalation policies, and post-incident reviews
  • Strong Terraform experience
  • Proficiency with PowerShell
  • Strong Azure experience and working knowledge of AWS
  • Azure DevOps CI/CD pipeline experience
  • Octopus Deploy experience
  • Comfort with Windows and Linux server environments
  • Familiarity with Slack
  • Experience with alert noise reduction, observability cost optimization, chaos engineering, SQL Server monitoring, FinTech compliance, reliability metrics, Python or Bash, or Jira is desirable

Benefits

  • Professional development budget
  • Flexible in-office days
  • Bi-weekly all-company meetings
  • Team offsites, team bonding activities, and happy hours
  • Health, retirement, family forming, and family support benefits
  • Employee giving match
  • Mobile phone stipend
  • R&R days
  • Wellness reimbursement and weekly onsite and virtual programming
  • Generous vacation policy
  • Parental leave and family planning benefits
  • Catered lunches and fully stocked kitchens
  • Competitive salary, bonuses, and equity

See also

要針對這個職缺調整履歷嗎?

目前無法檢查您與這個職缺的符合程度;請先將履歷加入個人檔案,下次即可查看。

A new version of freehire is available