Software Architect, Network System Validation

Summary

Architect the software systems that validate NVIDIA's high-speed networking (Ethernet, InfiniBand, NVLink, BlueField) powering AI clusters. Lead design of validation frameworks, chaos-injection orchestration, telemetry pipelines, and AI-driven test automation at scale.

NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. It’s a unique legacy of innovation that’s fueled by great technology—and amazing people. Today, we’re tapping into the unlimited potential of AI to define the next era of computing. An era in which our GPU acts as the brains of computers, robots, and self-driving cars that can understand the world. Doing what’s never been done before takes vision, innovation, and the world’s best talent. As an NVIDIAN, you’ll be immersed in a diverse, supportive environment where everyone is inspired to do their best work. Come join the team and see how you can make a lasting impact on the world.

Within NVIDIA, the Networking Business Unit (NBU) builds the high-speed interconnect — Ethernet, InfiniBand, NVLink, and BlueField DPUs — that switches thousands of GPUs into a single AI supercomputer, moving data at the scale and speed the most demanding workloads require.

NVIDIA is looking for a Software Architect to join our Network System Validation group. You will work on developing systems for validating advanced networking solutions across NVIDIA complex AI cluster environments. The group is a high-performance engineering force that treats validation as a first-class software problem. We build systems, frameworks, and benchmarks that prove our network's correctness and performance at scale. This is a senior, deeply hands-on role for a technology leader who can own the software architecture, development roadmap, mentor a team of high-performance engineers, and push NVIDIA's network to its speed-of-light limits. This role combines the design of validation methodologies with hands-on development of testing frameworks, automation infrastructure, and advanced debugging, analysis, and investigation tools for network performance and functionality at scale.

What you’ll be doing:

  • Design and implement network validation methodologies for system-level testing

  • Build the systems that orchestrate millions of concurrent IO operations, inject chaos at the infrastructure layer (latency, congestion, losses, and hardware failures), and expose the hardest-to-find system faults

  • Building the systems that capture, store and process TBs of telemetry data to produce autonomies root-cause-analysis and system level insights

  • Advance AI-driven approaches to test automation: intelligent scenario generation, LLM-augmented root-cause analysis, and autonomous validation pipelines

  • Establish engineering practices — design docs, production-grade code reviews, testing philosophy, and cross-team technical alignment

  • Produce clear, data-driven reports and insights

What we need to see:

  • B.Sc. / B.A. in Computer Science, Electrical Engineering, or equivalent experience

  • 12+ years of experience in developing large scale software projects

  • Strong distributed system experience and system-level debugging skills

  • Strong software skills: C/C++/Rust, Python (must), Bash

  • Experience building automation frameworks and validation tools

  • Experience building large-scale infrastructure platforms, internal developer platforms, or reliability engineering systems

  • Deep understanding of distributed systems: concurrency, consistency models, fault tolerance, and large-scale system behavior under stress

  • Proven track record leading complex technical initiatives from architecture through delivery

  • Ability to drive technical alignment across teams, communicate tradeoffs clearly, and make high-quality architectural decisions at speed

Ways to stand out from the crowd:

  • Background with RDMA / RoCE / AI networking (NCCL)

  • Experience with L2/L3/L4 networking protocols

  • Familiarity with NVIDIA networking solutions (ConnectX, SpecX, BlueField)

  • Background in performance analysis, Kubernetes, or cloud environments

  • Background in chaos engineering, fault injection, or simulation systems

We have some of the most forward-thinking and hardworking people working for us. If you're creative and autonomous, we want to hear from you!

NVIDIA is committed to fostering a diverse work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, disability status or any other characteristic protected by law.

See also

要針對這個職缺調整履歷嗎?

目前無法檢查您與這個職缺的符合程度;請先將履歷加入個人檔案,下次即可查看。

A new version of freehire is available