Cloud Hardware Development Engineer, AWS Hardware Engineering Services, Specialized Platforms and Servers

Summary

Design, automate, and debug AWS data center server hardware — building predictive failure detection, driving zero-touch fleet operations, and performing root-cause analysis across firmware, kernel, driver, thermal, and power layers.

Amazon Web Services (AWS) Hardware Engineering Services (HWEngS) owns the new product development (NPI) and operation of all AWS global infrastructure. In other words, we’re the people who keep the cloud running. We support all AWS data centers and all of the servers, storage, networking, power, and cooling equipment that ensure our customers have continual access to the innovation they rely on. We work on the most challenging problems, with thousands of variables impacting the supply chain, and we’re looking for talented people who want to help. You’ll join a diverse team of software, hardware, and network engineers, supply chain specialists, security experts, operations managers, and other vital roles. You’ll collaborate with people across AWS to help us deliver the highest standards for safety and security while providing seemingly infinite capacity at the lowest possible cost for our customers, and you’ll experience an inclusive culture that welcomes bold ideas and empowers you to own them to completion.

Cloud Hardware Development Engineer (CHDE): The Amazon Web Services (AWS) Hardware Engineering Services (HWEngS) Specialized Platforms and Servers team creates Enterprise rack solutions for Amazon’s innovative web services. We are seeking experienced CHDEs to own the fleet health, diagnostics, and automation of the Enterprise rack solutions:
1.) Designing and implementing predictive failure detection systems using telemetry, sensor data, error trends, and log correlation to identify hardware issues before they cause a customer impact.
2.) Driving toward zero-touch operations by building detection, diagnostics, and remediation of faults without human intervention
3.) Debugging complex system failures in time-sensitive settings personally diving deep when the problem demands it.
4.) Completing root cause analysis correlating across firmware, kernel, driver, thermal, power, and physical layers.

What you will do: As a member of the Specialized Platforms and Servers team, you’ll be responsible for collaborating with Elastic Cloud Compute (EC2) service teams and Data Center Operations to maintain fleet health in all the locations we have servers.

You will work closely with internal teams, suppliers, and external partners capturing lessons learned while operating the fleet to ensure next generation designs are of the highest quality, constantly looking for ways to improve your product performance, quality and cost.


Key job responsibilities
As a CHDE you will be responsible for scaling how we operate our massive existing & rapidly growing fleet. You will lead the integration and delivery of servers, support the development of automated monitoring, and failure analysis services to operate, debug, and scale our servers. You will work closely with other AWS software teams to tailor and operate servers solutions for the AWS environment. You will support launching our servers into production and operating our fleet of servers.

A day in the life
Your day to day responsibilities will be solving operational challenges to our existing fleet with the goal of improving the current customer experience as well as developing improved systems for future designs.

About the team
The team is comprised of CHDE's, System Development Engineers and Technical Program Managers, all with the common goal of delivering the best specialized server fleet possible to our customers.

Why AWS
Amazon Web Services (AWS) is the world’s most comprehensive and broadly adopted cloud platform. We pioneered cloud computing and never stopped innovating — that’s why customers from the most successful startups to Global 500 companies trust our robust suite of products and services to power their businesses.



Inclusive Team Culture
Here at AWS, it’s in our nature to learn and be curious. Our employee-led affinity groups foster a culture of inclusion that empower us to be proud of our differences. Ongoing events and learning experiences, including our Conversations on Race and Ethnicity (CORE) and AmazeCon (diversity) conferences, inspire us to never stop embracing our uniqueness.

Work/Life Balance
We value work-life harmony. Achieving success at work should never come at the expense of sacrifices at home, which is why we strive for flexibility as part of our working culture. When we feel supported in the workplace and at home, there’s nothing we can’t achieve in the cloud.

Mentorship and Career Growth
We’re continuously raising our performance bar as we strive to become Earth’s Best Employer. That’s why you’ll find endless knowledge-sharing, mentorship and other career-advancing resources here to help you develop into a better-rounded professional.

Diverse Experiences
Amazon values diverse experiences. Even if you do not meet all of the preferred qualifications and skills listed in the job description, we encourage candidates to apply. If your career is just starting, hasn’t followed a traditional path, or includes alternative experiences, don’t let it stop you from applying.

See also

要針對這個職缺調整履歷嗎?

目前無法檢查您與這個職缺的符合程度;請先將履歷加入個人檔案,下次即可查看。

A new version of freehire is available