Senior SRE Engineer
Maintain and operate AWS infrastructure for 24/7 stability and performance, monitor systems, handle incidents, automate operational tasks, optimize architecture, support high availability and backups, collaborate on system design and deployments, document operations, support internal IT needs, and participate in on-call rotation.
Responsibilities
- Maintain and operate AWS infrastructure to ensure 24/7 stability and performance
- Monitor systems using Zabbix and ELK, handle incidents, and develop custom scripts
- Analyze and resolve platform issues and optimize architecture and performance
- Support high availability, backup, and troubleshooting for applications and databases
- Collaborate with backend, product, and infrastructure teams on system design and deployment
- Document operations and incidents and support internal IT needs
- Participate in on-call rotation
Requirements
- 5+ years of Linux system administration experience
- Hands-on experience with AWS services including EC2, Lambda, Aurora, ElastiCache, CloudWatch, CloudFront, EKS, and IAM
- Proficiency in Bash, Python, and Golang scripting
- Experience with Kubernetes and infrastructure as code tools such as Terraform, Helm, and Kustomize
- Familiarity with Jenkins, GitHub Actions, Argo Workflows/CD, Airflow, and DAG development
- Knowledge of scalable system design, MongoDB, Kafka, load balancers, and message queues
- Understanding of information security best practices
- Strong problem-solving, communication, teamwork, and independent working skills