Infrastructure Engineer and SRE
Design and operate secure, scalable infrastructure for AI workloads, including untrusted code execution, parallel AI-agent runtimes, multi-cloud and customer VPC deployments, observability, and enterprise integrations.
Responsibilities
- Design systems for large-scale untrusted code execution and sandboxing
- Build massively parallel AI-agent runtime and scheduling systems
- Design multi-cloud and customer VPC deployment architecture
- Develop highly reliable distributed systems with strict security and data guarantees
- Implement observability across AI workflows and infrastructure
- Build enterprise-grade integration and metadata platforms
- Define how AI systems run safely in production
Requirements
- Strong background in distributed systems, scalability, multi-cloud architecture, and security
- Experience with operating systems, containers, Kubernetes, cloud networking, and event-driven runtimes such as Knative or KEDA
- Experience designing, building, and operating large-scale infrastructure
- Passion for building simple, elegant solutions to hard systems problems
- Strong zero-to-one mindset