Senior Site Reliability Engineer, Observability
This engineering-first role focuses on hands-on observability and reliability engineering across predominantly Windows-based Azure and AWS infrastructure supporting enterprise treasury customers. The role includes monitoring and alerting in New Relic, SLO/SLI and error-budget management, Terraform infrastructure as code, Azure DevOps pipelines, Incident.IO administration, incident response, and coaching engineering teams.
Responsibilities
- Design and implement monitoring, alerting, and dashboards in New Relic across Azure and AWS
- Write NRQL queries for troubleshooting, analysis, and reporting
- Define and implement SLOs, SLIs, and error budgets
- Lead alert noise reduction and signal quality engineering
- Optimize observability costs through log ingestion management and pipeline rules
- Develop and maintain Terraform infrastructure as code for monitoring and observability resources
- Establish IaC governance standards
- Author and troubleshoot Azure DevOps pipelines
- Administer Incident.IO, including alert routing and notification workflows
- Build incident management foundations, postmortem processes, on-call rotations, escalation policies, and response playbooks
- Track MTTR, MTTD, and incident frequency
- Respond to and debrief on production incidents
- Enable engineering teams through workshops, consultation, documentation, and training
Requirements
- 7+ years in Site Reliability Engineering, DevOps, or Platform Engineering focused on observability and production operations
- Hands-on engineering experience combined with coaching and mentoring ability
- Experience in Agile/Scrum environments
- Expert New Relic and NRQL experience
- Knowledge of structured logging, metrics collection, distributed tracing, dashboards, and alerts
- Expertise with SLOs, SLIs, and error budgets
- Experience with Incident.IO, PagerDuty, OpsGenie, or similar platforms
- Experience designing incident workflows, on-call rotations, escalation policies, and post-incident reviews
- Strong Terraform experience
- Proficiency with PowerShell
- Strong Azure experience and working knowledge of AWS
- Azure DevOps CI/CD pipeline experience
- Octopus Deploy experience
- Comfort with Windows and Linux server environments
- Familiarity with Slack
- Experience with alert noise reduction, observability cost optimization, chaos engineering, SQL Server monitoring, FinTech compliance, reliability metrics, Python or Bash, or Jira is desirable
Benefits
- Professional development budget
- Flexible in-office days
- Bi-weekly all-company meetings
- Team offsites, team bonding activities, and happy hours
- Health, retirement, family forming, and family support benefits
- Employee giving match
- Mobile phone stipend
- R&R days
- Wellness reimbursement and weekly onsite and virtual programming
- Generous vacation policy
- Parental leave and family planning benefits
- Catered lunches and fully stocked kitchens
- Competitive salary, bonuses, and equity