Finding the best job has never been easier

Limitless High-tech career opportunities - Expoint

Nvidia Senior HPC AI Networking Performance Research Analysis Engineer
United States, California
369732046

15.07.2025

US, CA, Santa Clara

time type: Full time

posted on: Posted 4 Days Ago

job requisition id

NVIDIA is seeking a Senior High Performance Computing (HPC) and AI Networking Performance Research and Analysis Engineer to join our Performance group. In this exciting role, you will profile and analyze AI workloads on large GPUs and CPUs scale clusters for distributed Deep Learning LLM training focused on collectives communication and networking. You will interact with many types of hardware and platforms, such as HCAs, Switches, CPUs, GPUs, and Systems. You will develop performance analysis tools and methodologies to dive deeply into the details and understand performance expectations, limitations, and bottlenecks.

What you'll be doing:

Exploring and researching AI workloads and DL models specifically tailored for large-scale deep learning LLM training on NVIDIA supercomputers and distributed systems focusing on high-performance networking and Nvidia Collective Communications Library (NCCL).
Benchmarking, Profiling, and Analyzing the performance to find bottlenecks and identify areas of improvement and optimizations, with a strong emphasis on networking aspects.
Implementing performance analysis tools.
Collaborating with many teams from hardware to software to provide performance analysis insights.
Defining performance test planning , setting performance expectations for new technologies and solutions, and working to reach the performance targets limits.

What we need to see:

B.Sc in Computer Science or Software Engineering or equivalent experience
5+ years of experience with high-performance Networking (RDMA, MPI, NCCL, Congestion Control Algorithms)
Demonstrated Performance Analysis skills and methodologies.
Experience with NVIDIA GPUs, CUDA library, deep learning frameworks like TensorFlow or PyTorch, combined with expertise in networking collective communication libraries (such as NCCL) and protocols (such as RoCE and RDMA).
Fast and self-learning capabilities with strong analytical and problem-solving skills.
Programming Languages: Python, Bash and C languages
Experience with Linux OS distros.
Great teammate with good communication and interpersonal skills

Ways to stand out from the crowd:

In-depth knowledge and experience with AI workloads and benchmarking for distributed LLM training.
Knowledge in CUDA, and NCCL libraries.
Knowledge in Congestion Control algorithms.
In-depth System knowledge and understanding (Intel / AMD / ARM CPUs, NVIDIA GPUs, HCA, Memory, PCI).
Strong Performance Analysis skills and methodologies using modern tools.

You will also be eligible for equity and .

Full job details

These jobs might be a good fit

Nvidia Senior AI Networking Performance Research Analysis Engineer China, Guangdong Province, Shenzhen

Nvidia Senior AI Networking Performance Research Analysis Engineer China, Shanghai

Nvidia Senior HPC Performance Engineer United States, California

Professional CV Builder tool from Expoint.

Get to the top of the "yes list" with a standout CV!

CREATE CV

Nvidia Senior HPC AI Networking Performance Research Analysis Engineer United States, California 369732046

Nvidia Senior AI Networking Performance Research Analysis Engineer China, Guangdong Province, Shenzhen

Nvidia Senior AI Networking Performance Research Analysis Engineer China, Shanghai

Nvidia Senior HPC Performance Engineer United States, California

Nvidia Senior HPC Performance Engineer United States, California

Nvidia Senior HPC AI Networking Performance Research Analysis Engineer
United States, California
369732046