SRE/DevOps Engineer- San Francisco, CA, the US
Ensure production reliability and runbooks, implement reliability guardrails on AWS, own deployment pipelines and GitHub code management, lead incident response and blameless postmortems, and operate monitoring, logging, and alerting systems.
Responsibilities
- Partner with Platform Engineering to implement AWS reliability guardrails meeting uptime and SLA requirements.
- Own deployment pipelines and code management practices via GitHub.
- Lead production incident troubleshooting and conduct blameless postmortems.
- Implement monitoring, logging, and alerting to detect and mitigate anomalies.
- Bridge US operations and international engineering hubs through bilingual technical communication.
Requirements
- Experience deploying and operating applications on AWS with mastery of GitHub workflows and actions.
- Strong monitoring, log management, and scripting skills for triage and troubleshooting.
- End-to-end ownership, transparency, responsiveness, and communication.
- High resilience and focus during high-stakes production incidents.
- Fluency in Mandarin and English, verbal and written.
Benefits
- Competitive packages aligned with California market standards
- Lead a dynamic and innovative team in a rapidly growing company
- Collaborative, inclusive environment where contributions are recognized and valued