Site Reliability Engineer
Own and improve the production environment for trading systems and exchange connectivity, focusing on performance, reliability, operability, monitoring, troubleshooting, tooling, operational risk, and technical operations mentorship.
Responsibilities
- Own the production environment
- Monitor and troubleshoot large-scale trading systems and exchange connectivity
- Build and maintain the DevOps toolkit
- Improve scalability and system performance using firm-wide metrics
- Analyze and troubleshoot complex system problems
- Coordinate changes and manage incidents
- Communicate technology changes with traders
- Reconcile trades and position breaks
- Assess operational risk of production changes
- Define and document processes and procedures
- Provide mentorship and cross-training
Requirements
- Degree in Computer Science, a related field, or equivalent professional experience
- At least 5+ years of relevant IT operations experience
- At least 3+ years of experience with Python and shell scripting
- Familiarity with C++
- Linux operating system knowledge
- Knowledge of network and system configuration
- Knowledge of kernel internals, scheduling, and performance tuning
- Networking knowledge including routing, multicast, LLDP, VLAN tagging, and Ethernet
- Ability to handle shared operational and periodic on-call duties
- Reliable and predictable availability