Cloud Operations & Support Engineer (L1 / L2)
Serve as the frontline for a public-cloud GPU platform by handling customer tickets, monitoring platform health, responding to alerts, troubleshooting infrastructure and service issues, escalating incidents, maintaining operational documentation, and reporting metrics.
Responsibilities
- Handle customer support tickets and requests
- Triage, troubleshoot, and resolve customer issues within SLA targets
- Monitor platform health and respond to NOC alerts
- Assess incident severity and initiate incident handling
- Provide customer-facing status updates
- Troubleshoot GPU instances, virtual machines, bare metal, networking, drivers, CUDA, storage, billing, and quota issues
- Escalate issues with complete context
- Maintain runbooks, knowledge bases, and canned responses
- Participate in a 24/7 follow-the-sun shift rotation
- Track and report ticket and alert metrics
Requirements
- 1+ years of experience in cloud or technical support, NOC, or IT operations
- Strong new graduates considered for L1
- Linux, networking, and cloud fundamentals
- Familiarity with GPU, CUDA, containers, virtual machines, bare metal, or storage
- Strong written English and customer communication
- Willingness to work follow-the-sun shifts
Benefits
- Attractive welfare benefits
- Training and mentoring
- Developmental opportunities
- Personal accountability, autonomy, fast growth, and learning opportunities
- Opportunity to contribute to new projects and processes