Sr Lead Infrastructure Engineer- Devops/AWS

We're looking for a hands-on DevOps / SRE engineer ready to take their career to new heights. Join the ranks of top talent at one of the world's most influential companies.

As a Lead DevOps / SRE Engineer - Vice President at JPMorgan Chase within the International Private Bank (IPB) Technology Artificial Intelligence and Machine Learning (AIML) Team, you will own the automation, reliability, and production operations of our agentic AI and machine learning products. As the team takes end-to-end ownership of the platforms it runs, you will build the CI/CD, observability, and incident-management practices that keep those services stable, secure, and performant across international markets.

This is a Vice President-level role and an integral part of the IPB Tech AIML team, reporting to the Head of AI, IPB Tech.

Job responsibilities

  • Owns and evolves the team's CI/CD pipelines, release automation, and deployment tooling
  • Establishes reliability practices (SLOs, error budgets, runbooks) and leads production incident response and post-incident review
  • Builds and operates observability across the team's AI/ML services (metrics, logging, tracing, alerting)
  • Automates infrastructure provisioning and configuration through infrastructure-as-code
  • Implements operational security, secrets management, and access controls to firm-wide standards
  • Partners with platform, data, and AI engineers to harden services for production and reduce dependency on external functions
  • Mentors junior engineers on DevOps and reliability practices and sets standards through review
  • Champions the firm's culture of diversity, Opportunity, inclusion, and respect

Required qualifications, capabilities, and skills

  • Formal training or certification on software engineering or systems concepts and applied experience
  • Advanced proficiency with infrastructure-as-code (e.g., Terraform) and scripting in Python and/or shell
  • Deep hands-on experience with CI/CD tooling and building release automation at scale
  • Strong experience with Kubernetes, containerisation, and cloud-native operations
  • Proven experience running production services: observability, on-call, incident response, and reliability engineering
  • Understanding of production security and change-management controls
  • Strong communication skills and the ability to set operational standards across a team
  • Formal SRE experience in a regulated or high-availability environment
  • Master's degree in Computer Science, Engineering, or a related technical field (or equivalent applied experience)

Preferred qualifications, capabilities, and skills

  • Experience operating ML / LLM workloads in production (MLOps, inference reliability, cost/performance management)
  • Experience within financial services technology
  • Familiarity with JPM-internal platform, cloud, and observability tooling for internal candidates

See also

要針對這個職缺調整履歷嗎?

目前無法檢查您與這個職缺的符合程度;請先將履歷加入個人檔案,下次即可查看。

A new version of freehire is available