Site Reliability Engineer
Summary
Site Reliability Engineer building reliable, automated, and observable production systems for mission-critical applications. Core stack spans Linux, Docker, Kubernetes, CI/CD, Terraform/Ansible IaC, and Grafana/Prometheus/ELK on AWS/Azure/GCP.
We are looking for a Site Reliability Engineer (SRE) with experience in platform engineering, DevOps, and production operations. The role involves building reliable systems, automating infrastructure, and ensuring observability across mission-critical applications. You will work closely with development, infrastructure, and operations teams to design scalable solutions, improve system resilience, and define operational best practices.
Key Responsibilities
System reliability: Ensure high availability, performance, and resilience of production systems.
Infrastructure automation: Build and maintain CI/CD pipelines, automate deployments, and manage containerized workloads.
Observability & monitoring: Implement logging, metrics, tracing, alerting, and dashboards using modern monitoring tools.
Incident response: Integrate systems with enterprise monitoring, SIEM, and incident management workflows; participate in on-call rotations.
Operational excellence: Define and implement runbooks, escalation paths, and production support models.
Collaboration: Work with cross-functional teams in Agile environments to deliver reliable and secure systems.
Continuous improvement: Conduct root cause analysis, implement permanent fixes, and drive automation initiatives to reduce manual effort.
Skills Required
5+ years of experience as an SRE, DevOps Engineer, Platform Engineer, Infrastructure Engineer, or Production Engineer.
Solid knowledge of Linux administration, Shell scripting, Git, CI/CD, Docker, and Kubernetes.
Hands-on experience with observability platforms (Grafana, Prometheus, ELK, CloudWatch, Azure Monitor, etc.).
Experience integrating systems with enterprise monitoring, alerting, SIEM, or incident response workflows.
Proven ability to define and implement runbooks, operational procedures, escalation paths, and production support models.
Familiarity with cloud platforms (AWS, Azure, GCP) and infrastructure-as-code tools (Terraform, Ansible).
Excellent problem-solving skills, ability to troubleshoot complex systems, and experience in Agile/DevOps practices.
Excellent communication skills and ability to collaborate with cross-functional stakeholders.
EA License : 02C3423EA Personnel : R22108699