Senior MLOps Engineer, LLMOps
Build and maintain infrastructure and pipelines for production AI systems, including CI/CD workflows, model versioning and approvals, compliance, observability, scalable model serving, offline and online evaluation, monitoring, and reproducible research environments.
Responsibilities
- Build reusable CI/CD workflows for model training, evaluation, and deployment
- Automate model versioning, approval workflows, and compliance checks
- Build modular and scalable AI infrastructure including vector databases, feature stores, model registries, and observability tooling
- Embed AI models and agents into real-time applications and workflows
- Evaluate and integrate state-of-the-art AI tools
- Drive AI reliability, governance, compliance, security, and uptime
- Ensure data accuracy, consistency, and reliability for training and inference
- Deploy infrastructure for offline and online evaluation, regression testing, cost monitoring, and human-in-the-loop workflows
- Provide sandboxes, dashboards, and reproducible environments for researchers
Requirements
- Write high-quality, maintainable software primarily in Python
- Experience with containerization and orchestration such as Docker and Kubernetes
- Experience with infrastructure-as-code and deployment tooling such as Terraform and CI/CD pipelines
- Experience with monitoring and logging frameworks such as Datadog, Prometheus, and OpenTelemetry
- Implement MLOps best practices including model versioning, rollback strategies, automated evaluation, and drift detection
- Experience with scalable model and agent serving infrastructure such as vLLM, Triton, and BentoML
- Experience deploying and maintaining LLM and agentic workflows in production
- Monitor cost, latency, and performance and capture traces for analysis and debugging
- Demonstrate strong ownership and pragmatism while balancing infrastructure elegance with iterative delivery
Benefits
- Equity plan eligibility