Expoint - all jobs in one place

The point where experts and best companies meet

Limitless High-tech career opportunities - Expoint

Nvidia Principal Solution Engineer Large Scale ML Profiling Services 
United States, California 
736803890

01.12.2024

What you’ll be doing:

You will plan and drive scalable profiling initiatives for ML customers. Prototype, design, and develop robust profiling solutions for large-scale ML use cases. Design actionable performance metrics for ML engineers to optimize workloads on CPU and GPU. Develop and implement deployment plans for software delivery and test profiling services in datacenter environments. Identify user friction points and emerging needs. Provide technical leadership to help strategies and status reports to senior leadership. Stay updated on the latest techniques and frameworks for AI training, deployment, and performance optimization.

What we need to see:

  • Ability to work in a multifaceted, fast-paced environment.

  • Experience with customer support, focusing on product and feature delivery

  • 15+ years of proven experience in system design, performance analysis, and shipping production software.

  • BS or higher degree in CS/EE/CE or equivalent experience.

  • Strong interpersonal skills.

  • Strong verbal and written communication skills

  • Experience deploying software in microservices distributed environments.

  • Strong software development experience in Python and C++.

  • Experience collaborating with different teams across the company.

Ways to stand out from the crowd:

  • Experience with performance analysis of AI training/inference applications.

  • Experience in building continuous profiling systems for GPU data centers.

  • Knowledge of GPU architecture and programming (CUDA, OpenCL).

  • Ability to tackle ambiguous situations and make them tractable.

You will also be eligible for equity and .