DevOps SRE
This role owns the reliability, stability, availability, security, and performance of production systems. The engineer will operate a modern cloud-native stack, support monitoring and incident response, manage infrastructure and deployments, and participate in on-call rotations.
Responsibilities
- Own the reliability, availability, and performance of production environments.
- Operate production Kubernetes on EKS, including cluster upgrades and Helm deployments.
- Manage scaling and capacity with KEDA, Karpenter, and HPA.
- Manage AWS cloud environments, including EC2, Lambda, AWS Batch, ElastiCache, and RDS.
- Evolve infrastructure as code with Terraform and Helm using security best practices.
- Support GitLab CI/CD pipelines and improve deployment stability.
- Design observability systems with Prometheus, Grafana, and EFK.
- Troubleshoot TLS, load balancing, VPC, NAT, and VPN issues.
- Support compliance initiatives and respond to security incidents.
- Use AI-powered tools for automation and productivity.
- Lead end-to-end incident response through troubleshooting, mitigation, and resolution.
- Perform deep-dive root-cause analysis.
- Participate in on-call rotations.
Requirements
- 3+ years of hands-on DevOps or SRE experience.
- Production experience with Docker and Kubernetes.
- Knowledge of AWS, including EKS, EC2, Organizations, RDS, S3, CloudWatch, Lambda, and DynamoDB.
- Experience with monitoring, logging, and alerting systems.
- Proficiency with Terraform, Helm, and GitLab CI or similar tools.
- Troubleshooting skills across infrastructure, CI/CD, and networking.
- Bash and Python scripting experience.
- Willingness to participate in on-call rotations.
- Familiarity with pub/sub systems such as SQS or Kafka.
- Experience with Redis, Airflow, Databricks, Spark, or EMR is a plus.
- Experience with GitOps workflows and advanced Git usage is a plus.
- Experience supporting Postgres, Snowflake, or ClickHouse is a plus.