Senior Site Reliability Engineer
Ensure the reliability, scalability, and performance of production systems while bridging development and operations. Own production health, deployment practices, monitoring, incident response, and cross-border technical alignment.
Responsibilities
- Implement AWS reliability guardrails to meet uptime and SLA requirements.
- Own deployment pipelines and repository management practices through GitHub.
- Lead production incident troubleshooting and conduct blameless postmortems.
- Implement monitoring, logging, and alerting to detect and mitigate anomalies.
- Bridge US operations and international engineering teams through bilingual technical collaboration.
Requirements
- Experience deploying and running applications on AWS with advanced Git workflows and GitHub Actions.
- Proficiency in monitoring tools, log management, and scripting for troubleshooting.
- Strong ownership, transparency, communication, and end-to-end responsibility for production health.
- High psychological resilience and composure during high-stakes incidents.
- Fluency in Mandarin and English, both verbal and written.
Benefits
- Competitive packages aligned with California market standards
- Lead a dynamic and innovative team in a rapidly growing company
- Collaborative, inclusive environment where contributions are recognized and valued