Expoint - all jobs in one place

Finding the best job has never been easier

Limitless High-tech career opportunities - Expoint

Nvidia Senior Site Reliability Engineer Data Science ML Platforms 
United States, Texas 
774307007

31.07.2024

What you’ll be doing:

  • Develop software solutions to ensure reliability and operability of large-scale systems supporting machine-critical use cases.

  • Gain a deep understanding of our system operations, scalability, interactions, and failures to identify improvement opportunities and risks.

  • Create tools and automation to reduce operational overhead and eliminate manual tasks.

  • Establish frameworks, processes, and standard methodologies to enhance operational maturity, team efficiency, and accelerate innovation.

  • Define meaningful and actionable reliability metrics to track and improve system and service reliability.

  • Oversee capacity and performance management to facilitate infrastructure scaling across public and private clouds globally.

  • Build tools to improve our service observability for faster issue resolution.

  • Practice sustainable incident response and blameless postmortems

What we need to see:

  • Minimum of 5-8 years of experience in SRE, Cloud platforms, or DevOps with large-scale microservices in production environments.

  • Master's or Bachelor's degree in Computer Science or Electrical Engineering or CE or equivalent experience.

  • Strong understanding of SRE principles, including error budgets, SLOs, and SLAs.

  • Proficiency in incident, change, and problem management processes.

  • Skilled in problem-solving, root cause analysis, and optimization.

  • Experience with streaming data infrastructure services, such as Kafka and Spark.

  • Expertise in building and operating large-scale observability platforms for monitoring and logging (e.g., ELK, Prometheus).

  • Proficiency in programming languages such as Python, Go, Perl, or Ruby.

  • Hands-on experience with scaling distributed systems in public, private, or hybrid cloud environments.

  • Experience in deploying, supporting, and supervising services, platforms, and application stacks.

Ways to stand out from the crowd:

  • Experience operating large-scale distributed systems with strong SLAs.

  • Excellent coding skills in Python and Go and extensive experience in operating data platforms.

  • Knowledge of CI/CD systems, such as Jenkins and GitHub Actions.

  • Familiarity with Infrastructure as Code (IaC) methodologies and tools.

  • Excellent interpersonal skills for identifying and communicating data-driven insights.

You will also be eligible for equity and .