網站可靠性工程:3,072 個職缺
瀏覽與 網站可靠性工程 相關的開放職缺。
Software Engineering Manager-Site Reliability Engineering Center
Position Overview At PNC, our people are our greatest differentiator and competitive advantage in the markets we serve. We are all united in delivering the best experience for our customers. We work together each day…
Senior Site Reliability Engineer
Principal Staff SRE at NVIDIA leading core IT infrastructure (DNS, NTP/PTP, DHCP, LDAP) across on-prem and cloud at global scale, with hands-on work in Linux kernel, eBPF/XDP, SR-IOV/DPU, Go/Python, Terraform, and large-scale bare-metal environments.
Site Reliability Engineer
Site Reliability Engineer building reliable, automated, and observable production systems for mission-critical applications. Core stack spans Linux, Docker, Kubernetes, CI/CD, Terraform/Ansible IaC, and Grafana/Prometheus/ELK on AWS/Azure/GCP.
Sr Mgr, Site Reliability Engineer (SRE)
Job Posting Title: Sr Mgr, Site Reliability Engineer (SRE) Req ID: 10145210 Job Description: At Disney, we’re storytellers. We make the impossible possible. The Walt Disney Company (TWDC) is a world-class entertainment…
Sr. Site Reliability Engineer
Calling all innovators - find your future at Fiserv. We're Fiserv, a global leader in Fintech and payments, and we move money and information in a way that moves the world. We connect financial institutions,…
Senior Staff SRE – Compute Platform
NVIDIA is seeking a Senior Staff SRE to build and operate reliable, scalable compute platforms that support global engineering workloads. This role spans Kubernetes, KubeVirt, bare-metal infrastructure, automation,…
Senior Site Reliability Engineering, Storage
NVIDIA has been redefining computer graphics, PC gaming, and accelerated computing for more than 25 years. It’s an outstanding legacy of innovation that’s fueled by phenomenal technology – and amazing people. Today,…
Senior Site Reliability Engineering - Storage
NVIDIA has been redefining computer graphics, PC gaming, and accelerated computing for more than 25 years. It’s an outstanding legacy of innovation that’s fueled by phenomenal technology – and amazing people. Today,…
Senior Network Site Reliability Engineer
We are seeking a highly skilled and experienced Staff Network Site Reliability Engineer (SRE) to join our Enterprise Network Operations and SRE team. In this role, you will be pivotal in implementing our vision for a…
Data Site Reliability Engineer (SRE)
Type of Requisition: Regular Clearance Level Must Currently Possess: None Clearance Level Must Be Able to Obtain: None Public Trust/Other Required: BI Full 6C (T4) Job Family: IT Infrastructure and Operations Job…
Senior Associate, Dev Ops Engineer, SRE and Governance, Group Technology
Senior DevOps/SRE role at DBS Bank in Singapore governing production releases, deployments, and privileged access across banking systems. Core stack spans Linux, Oracle/PostgreSQL, CyberArk PAM, OpenShift, and CI/CD automation in a regulated 24x7 environment.
AVP, Platform SRE Engineer, SRE&Governance, Group Technology
AVP-level Platform SRE Engineer at DBS Bank ensuring high availability of enterprise observability and analytics platforms. Day-to-day work centers on AppDynamics/ELK/Grafana/OpenTelemetry monitoring, automation scripting in Python/Shell, performance troubleshooting, and CI/CD integration.
Site Reliability Engineer (Splunk, Python, OCI, Dynatrace, RCA, Terraform, Ansible )
Responsibilities Manage and maintain highly available, scalable, and reliable production environments. Lead incident management, root cause analysis (RCA), and problem management activities to ensure service stability.…
Golang Engineer - SRE Engineering Productivity
Build and operate the CI/CD, testing, and infrastructure tooling that other Arista development teams rely on. Day-to-day work spans Go development, workflow automation, and observability using Ansible, Kubernetes, Jenkins, Grafana, Spinnaker, MySQL, ElasticSearch, GCP, and Varnish.
SRE (Site Reliability Engineer)
SRE at ROUTE06 ensuring reliability of their enterprise commerce API platform 'Plain', which powers mall-type EC, OMO stores, and marketplaces. Day-to-day work spans infrastructure setup, monitoring, operation automation, performance tuning, and CI/CD on AWS/GCP, with Docker and IaC.
SRE-инженер
Design fault-tolerant architecture and observability (metrics, logs, tracing) for corporate IT infrastructure; investigate incidents, run post-mortems, set up meaningful alerting and reliability processes. Core stack: Prometheus, Grafana, Loki, Tempo, Docker, Kubernetes on Linux.