Sr Network Engineer, Reliability and Observability (USA)
Summary
Senior network engineer combining production datacenter networking with software engineering to build reliability and observability automation for a high-performance AI/accelerated-computing infrastructure platform. Day-to-day work centers on telemetry, incident response, and tooling in Golang (with Python or Rust).
Sr Network Engineer, Reliability and Observability
Location: Remote, United States or London, United Kingdom
US Location Preference: New York, San Francisco, Austin, Seattle, or Phoenix
Our client is scaling high-performance infrastructure that supports advanced artificial intelligence and accelerated computing workloads. They are seeking a Sr Network Engineer, Reliability and Observability who can combine deep production networking expertise with software engineering and a data-driven approach to infrastructure reliability. This engineer will help ensure complex datacenter fabrics remain resilient as the environment grows, while building the automation, telemetry, and engineering processes needed to identify problems earlier and prevent repeat failures.
This Role Offers:
- Exposure to sophisticated datacenter fabrics, high-performance networking, optics, and physical infrastructure.
- Opportunity to build software and automation that improves network operations rather than relying solely on manual processes.
- Competitive compensation with equity participation and a comprehensive benefits package.
Focus:
- Drive network reliability programs by turning telemetry, operational trends, incident history, and infrastructure data into measurable improvements in availability and performance.
- Design automated data workflows, monitoring systems, and operational tooling that provide continuous visibility into network health, service performance, and recurring failure patterns.
- Support high-performance data center networking environments by isolating and correcting issues involving traffic flow, network design, system configurations, switching infrastructure, and physical connections.
- Take technical ownership during complex production incidents, systematically isolate failure domains, coordinate response efforts, and carry issues through remediation and follow-up improvements.
- Build scalable infrastructure and reliability software primarily in Golang, with Python or Rust used for supporting tools and automation, while applying disciplined development and testing practices.
- Collaborate with deployment, datacenter operations, hardware, logistics, and software teams to improve infrastructure lifecycle processes, remediation workflows, and operational readiness.
- Serve as a technical resource across several critical infrastructure domains, bringing specialized knowledge in areas such as network communications, high-speed connectivity, fiber systems, data transport, or power technologies.
- Translate ambiguous infrastructure challenges into defined objectives, Jira initiatives, development pipelines, working code, documentation, and sustainable operational improvements.
Skill Set:
- 5+ years of professional experience supporting network infrastructure, with at least 3 years focused on hands-on operational responsibilities within live production or high-performance computing environments.
- Extensive experience supporting sophisticated data center networks, including routed and switched architectures, overlay technologies, dynamic routing, large-scale fabric designs, configuration troubleshooting, and connectivity issues across both logical and physical infrastructure.
- Strong software engineering background, including experience with ITIL, Agile/xP, and TDD methodologies and development of hyperscale platforms using Golang with Python or Rust tooling.
- Proven ability to lead incident response, troubleshoot complex infrastructure failures methodically, communicate effectively during outages, and maintain ownership through resolution.
- Demonstrated technical depth across multiple infrastructure disciplines, with strong expertise in at least two areas such as network protocols, fiber technologies, transport systems, high-performance connectivity, or power infrastructure.
- Proven ability to independently turn broad technical goals into structured plans and carry them through development, validation, implementation, and final delivery.
- Strong operational judgment with the ability to prioritize effectively and make sound technical decisions when working with incomplete information or time-sensitive production issues.
- Effective communication and documentation skills with the ability to convert technical findings and incident lessons into repeatable engineering practices.
- Additional value placed on experience supporting high-performance computing infrastructure, advanced network technologies, monitoring and analytics platforms, data-driven troubleshooting, or hands-on data center operations.
About Blue Signal:
Blue Signal is an award-winning, executive search firm specializing in various specialties. Our recruiters have a proven track record of placing top-tier talent across industry verticals, with deep expertise in numerous professional services. Learn more at bit.ly/46Gs4yS