Expoint – all jobs in one place
מציאת משרת הייטק בחברות הטובות ביותר מעולם לא הייתה קלה יותר
Limitless High-tech career opportunities - Expoint

Nvidia Senior Site Reliability Engineer AI Infrastructure 
United States, Texas 
52266324

Today
US, CA, Santa Clara
US, CA, Remote
US, OR, Remote
time type
Full time
posted on
Posted 8 Days Ago
job requisition id

What You Will Be Doing:

  • Develop and maintain large-scale systems supporting critical use cases for AI Infrastructure, driving reliability, operability, and scalability across global public and private clouds.

  • Implement SRE fundamentals, including incident management, monitoring, and performance optimization, while designing automation tools to reduce manual processes and operational overhead.

  • Build tools and frameworks to improve observability, define actionable reliability metrics, and enable fast issue resolution, driving continuous improvement in system performance.

  • Establish frameworks for operational maturity, lead sustainable incident response protocols, and conduct blameless postmortems to improve team efficiency and system resilience.

  • Work with engineering teams to deliver innovative solutions, mentor peers, uphold high standards for code and infrastructure, and contribute to hiring for a diverse, high-performing team.

What We Need to See:

  • Degree in Computer Science or related field, or equivalent experience with 12+ years in Software Development, SRE, or Production Engineering.

  • Proficiency in Python and at least one other language (C/C++, Go, Perl, Ruby).

  • Expertise in systems engineering within Linux or Windows environments and cloud platforms (AWS, OCI, Azure, GCP).

  • Strong understanding of SRE principles, including error budgets, SLOs, SLAs, and Infrastructure as Code tools (e.g., Terraform CDK).

  • Hands-on experience with observability platforms (e.g., ELK, Prometheus, Loki) and CI/CD systems (e.g., GitLab).

  • Strong communication skills with the ability to convey technical concepts effectively to diverse audiences.

  • Commitment to fostering a culture of diversity, curiosity, and continuous improvement.

Ways to stand out from the crowd:

  • Experience in AI training, inferencing, and data infrastructure services.

  • Proficiency in deep learning frameworks like PyTorch, TensorFlow, JAX, and Ray.

  • A strong background in hardware health monitoring and system reliability.

  • Hands-on expertise in operating and scaling distributed systems with stringent SLAs, ensuring high availability and performance.

  • Proven experience in incident, change, and problem management processes, fostering continuous improvement in sophisticated environments.

You will also be eligible for equity and .