Expoint – all jobs in one place
The point where experts and best companies meet
Limitless High-tech career opportunities - Expoint

Nvidia Senior AI Infrastructure Engineer Cloud Partnerships - DGX 
United States, Texas 
307198966

Today
US, CA, Santa Clara
US, Remote
time type
Full time
posted on
Posted 3 Days Ago
job requisition id

What you’ll be doing:

  • Architect unified systems for integrating infrastructure provider maintenance events into NVIDIA engineering systems

  • Drive the adoption of operational excellence best practices across all infrastructure providers, partnering with SRE, infra, product, and security teams

  • Define and operationalize governance models for engineering support engagements, infrastructure maintenance lifecycles, and incident escalation paths

  • Measure provider availability against projected maintenance schedules using Service Level Objectives (SLOs)

  • Collaborate with AI/ML teams to integrate intelligent automation into maintenance workflows, such as projecting job capacity impact based on scheduled resource availability and suggesting infrastructure reallocations for high-profile initiatives

  • Develop a long-term roadmap to guide infrastructure providers in progressively adopting best practices for reliability and production hygiene across existing and new product introductions

What we need to see:

  • 8+ years of experience in infrastructure architecture, cloud native, or large-scale platform/reliability roles

  • Bachelor's degree or equivalent experience

  • Experience designing scalable, maintainable backend systems and writing clear design documentation

  • Strong understanding of multiple cloud infrastructure provider resource offerings

  • Demonstrated experience in normalizing and unifying diverse data sources from a variety of systems into broadly applicable schemas, enabling efficient querying and analysis

  • Proven ability to lead and influence cross-functional technical initiatives at scale across vendors and external partners, especially in reliability or platform domains

  • Demonstrated ability to design and implement maintainable APIs for internal and external customers

  • Proficiency in Kubernetes administration, modern CI/CD techniques and Infrastructure as Code (IaC)

  • Experience building resilient production systems using Golang, Python or Ruby

Ways to stand out from the crowd:

  • Proven experience in operating or architecting production systems across multiple cloud and bare-metal infrastructure providers.

  • Proficiency in data science tools such as Spark, Delta Lake, and Databricks.

  • Hands-on production experience with workflow engines like Temporal, Argo, or similar platforms for durable workflow execution.

You will also be eligible for equity and .