Founding Infrastructure Engineer
Technical owner of the operational backbone for a public AI inference utility, responsible for reliability, observability, routing transparency, production ML serving infrastructure, and coordination across cloud and HPC partners.
Responsibilities
- Harden the platform for major launches
- Perform load testing and build fallback routing
- Set up monitoring and end-to-end observability
- Ship downtime warnings and fallback behavior
- Implement routing transparency and endpoint provenance
- Improve public endpoint performance
- Integrate MCP and programmatic infrastructure interfaces
- Operate production inference and ML serving infrastructure
- Coordinate with cloud providers and HPC centers
- Orchestrate and integrate open-source stacks
Requirements
- Significant experience operating production inference or ML serving infrastructure
- Strong distributed systems and SRE instincts
- Experience with observability, incident response, fallback design, and capacity planning
- Comfort working with cloud providers, sovereign HPC centers, and institutional IT
- Experience orchestrating multiple stacks and open-source projects
- Maintainer and integrator experience
- Ability to work autonomously in a small team
- Ability to travel occasionally for team workshops