SRE Manager
Lead and grow the SRE team to maintain highly available, scalable, and reliable production systems. Own AWS infrastructure operations, monitoring, security, cost optimization, incident response, compliance, observability, CI/CD, infrastructure as code, GitOps, Kubernetes, and operational improvements.
Responsibilities
- Lead and manage the SRE team.
- Own AWS cloud infrastructure operations, monitoring, security, resource management, and cost optimization.
- Lead incident management, troubleshooting, root-cause analysis, and post-incident improvements.
- Ensure infrastructure and operational processes meet security, audit, and regulatory requirements.
- Drive observability, alerting, SLA/SLO/SLI, capacity planning, disaster recovery, and high availability practices.
- Improve system performance, reliability, and operational efficiency through automation and architecture optimization.
- Build and maintain CI/CD, IaC, and GitOps workflows.
- Manage Kubernetes, EKS, and containerized infrastructure.
- Collaborate with Backend, Data, Security, and Product teams.
- Build and improve monitoring and observability platforms.
- Mentor team members and maintain operational documentation and incident processes.
Requirements
- 8+ years of Linux system administration and large-scale infrastructure experience.
- 2+ years of team management or Tech Lead experience.
- Experience operating high-traffic, high-availability cloud platforms in a 24/7 environment.
- Strong AWS experience, including EC2, API Gateway, AppSync, VPC, IAM, Lambda, Aurora, ElastiCache, CloudFront, CloudWatch, EKS, SNS, Parameter Store, and Secrets Manager.
- Strong Kubernetes and container infrastructure experience, including EKS administration and troubleshooting.
- Experience with Terraform, Helm, and Kustomize.
- Experience with Jenkins, GitHub Actions, Argo Workflow, and ArgoCD.
- Experience with Grafana, ELK, Zabbix, and Nagios.
- Experience managing MongoDB, Kafka, load balancers, and high-availability architectures.
- Strong understanding of SRE and DevOps practices.
- Proficiency in Bash, Python, or Golang.
- Knowledge of cloud security, infrastructure security, and technical risk management.
- Strong communication, collaboration, and problem-solving skills.
- FinTech, Crypto, MAS TRM, and ISO 27001 experience are advantageous.