Expoint – all jobs in one place
Finding the best job has never been easier
Limitless High-tech career opportunities - Expoint

Nvidia DL System Software Engineer - AI Platform 
Canada, Ontario, Old Toronto 
817372857

09.09.2025
Canada, Toronto
time type
Full time
posted on
Posted 6 Days Ago
job requisition id

What you will be doing:

  • Taking part in the development of the NVIDIA's AI platform for training, fine-tuning and serving latest and greatest AI models with the best performance and efficiency.

  • Designing and building solutions for scheduling large scale AI training and inference workloads on GPU clusters over many cloud infrastructure.

  • Exploring and finding solution for open problems like industry-scale resource management, GPU scheduling, performance prediction, and live workload migration.

  • Work with and contribute to adjacent teams like TensorRT/Dynamo inference engine, ML compiler, KAI/Grove scheduler, Lepton cloud etc.

What we need to see:

  • Bachelor's degree or equivalent experience in Computer Science, Computer Engineering, relevant technical field.

  • 5+ years of experience.

  • Experience building large scale systems from scratch. Prior experience in container-based deployment systems like Kubernetes is beneficial.

  • Strong coding skills in programming languages like Python, Go, Rust and/or C/C++.

  • Solid foundation in other computer science and computer engineering topics: algorithms and data structures, operating systems, computer architecture, etc. Strong understanding of AI and related technologies is a huge plus.

  • Most importantly, ability to quickly grasp new concepts and thrive in evolving situations.

Ways to stand out from the crowd:

  • Graduate-level education or relevant practical background, particularly in research, is beneficial.

  • Practical experience in building and optimizing AI applications is highly desired.

  • Proficiency in container software such as containerd, CRI-O, Linux namespace, CRIU, and NVIDIA GPU technology such as CUDA graphs, Driver/runtime is greatly advantageous.

You will also be eligible for equity and .

Applications for this job will be accepted at least until September 6, 2025.