Site Reliability Engineer | Weekend Warrior
Manage the real-time production trading environment by monitoring and troubleshooting large-scale trading systems and exchange connectivity. Build reliability tools, improve scalability and performance, coordinate deployments, manage incidents, reconcile trades, assess operational risk, document procedures, and mentor technical operations engineers. Work a four-day schedule that includes one weekend day.
Responsibilities
- Own the production environment and drive performance, reliability, and operability improvements
- Monitor and troubleshoot large-scale trading systems and exchange connectivity
- Build and maintain site reliability tools for configuration management, process management, deployment, monitoring, data collection, and analysis
- Use firm-wide metrics to improve scalability and system performance
- Coordinate technology changes and deployments with traders, Risk Management, and Operational Trading Support teams
- Analyze and troubleshoot complex system problems
- Reconcile trades and position breaks with the Clearing team
- Assess and manage operational risk when deploying production changes
- Define and document processes and procedures
- Mentor and cross-train technical operations engineers
Requirements
- Degree in Computer Science, a related field, or equivalent professional experience
- 5+ years of relevant experience in IT operations, DevOps, SRE, Linux Systems Engineering, or Network Engineering
- 3+ years of experience with Python and shell scripting
- Linux operating system knowledge
- Networking knowledge including routing, multicast, LLDP, VLANs, and Ethernet
- Ability to handle shared operational and periodic on-call duties
- C++ experience is a plus
Benefits
- Private Medical, Vision and Dental Insurance
- Travel Medical Insurance
- Group Pension Scheme
- Group Life Assurance and Income Protection Schemes
- Paid Parental Leave
- Parking and Commuter Benefits