Middle SRE Engineer
Summary
RedCore is hiring a remote Middle SRE Engineer to run production infrastructure and reliability practices, building Terraform/Helm/ArgoCD modules, supporting incidents and DRP, and improving observability across Kubernetes and AWS stacks using GitLab CI and Grafana.
RedCore is an international business group that creates technological solutions for digital markets. Our products and services cover fintech, marketing, e-commerce, customer service, communications, and regulatory technologies.
We are looking for a Middle SRE Engineer!
Requirements:
- Experience with Kubernetes and Helm;
- Familiarity with infrastructure management using Terraform and Ansible;
- Awareness of AWS services such as VPC, IAM, S3, EC2, RDS, SM, SSM, EKS, ECR;
- Ability to read and understand Golang code;
- Experience with operating systems using systemd and bash (POSIX);
- Exposure to handling CI/CD pipelines using tools like GitLab CI and ArgoCD or similar tools;
- Basic knowledge of monitoring stacks, with a preference for Grafana.
Will be a plus:
- Understanding of large-scale systems management and design;
- Experience in building software applications.
Responsibilities:
Incident management
- Support and resolve incidents in the production environment, including participation in postmortems and incident cause analysis
- Participate in the development of DRP practices: document scenarios, conduct training
- Participate in the setup and diagnosis of replication and backups according to existing policies
- Maintain the relevance and reliability of the backup and recovery scheme
Infrastructure management
- Support for available infrastructure
- Optimization of infrastructure and systems productivity
- Development and scaling of the system
- Setting up new "instances" of infrastructure: environments, clusters, instances, services, etc.
- Contribute to the company's Terraform, Helm, and ArgoCD modules
- Develop and implement metrics and dashboards to improve service observability under existing policies
- Integrate and enhance tracking of services and products uptime metrics
- Analyze existing processes, develop, and implement automated solutions
Uptime tracking and Reliability practices
- Contribute to uptime tracking strategy development
- Set up monitoring and alerting for releases, infrastructure, and critical systems
- Review and provide feedback on internal infrastructure, applications, and modules
- Implementation of new tools, approaches, and practices
- Daily team communication
Our benefits to you:
🍀 An exciting and challenging job in a fast-growing business group, the opportunity to be part of a multicultural team of top professionals in Development, Architecture, Management, Operations, Marketing, Legal, Finance, and more
🤝🏻 Great working atmosphere with passionate experts and leaders, sharing a friendly culture and a success-driven mindset is guaranteed
🧑🏻💻 Modern corporate equipment based on macOS or Windows, and additional equipment is provided
🏖️ Paid vacations, sick leave, personal events days, days off
💵 Referral program — enjoy cooperation with your colleagues and get a bonus
📚 Educational programs: regular internal training sessions, compensation for external education, attendance of specialized global conferences
🎯 Rewards program for mentoring and coaching colleagues
🗣️ Free internal English courses
✈️ In-house Travel Service
🦄 Multiple internal activities: online platform for employees with quests, gamification, presents and news, clubs for movie/book/pets lovers, and more
🎳 Other benefits could be added based on your location