JobsSenior Site Reliability Engineer - HPC
NVIDIA logo

Senior Site Reliability Engineer - HPC

NVIDIA

Location

Santa Clara, CA, Austin, TX, Durham, NC

Type

Full-time

Posted

6/16/2026

Compensation

$152,000 - $287,500 per year

Undergraduate with 5+ Years of Experience
Approval 99.2%·Filings 1,781·New hires 873·
👑 Elite Sponsor
·FY 2025

Job description

NVIDIA is seeking a Senior Site Reliability Engineer (SRE) to join the Compute Farm team, focusing on building the next generation of their global services platform. The role involves ensuring the reliability and performance of critical systems while leveraging AI technologies to solve complex problems. The ideal candidate will have a strong background in supporting large-scale HPC clusters and will work in a fast-paced environment to deliver innovative solutions. This position offers the opportunity to have a significant impact on the future of computing.

Requirements

  • B.S. degree in Computer Science or related technical field or equivalent experience with 5+ years of professional experience building and supporting critical services.
  • Experience supporting large-scale HPC clusters using Slurm, LSF, or Kubernetes clusters.
  • Proficiency in modern CI/CD techniques and Infrastructure as Code (IaC) for managing services.
  • Strong experience in crafting large-scale infrastructure platforms for automated host lifecycle management and fleet reliability.
  • 5+ years of coding/scripting experience in at least two high-level programming languages such as Python, Go, Perl, or Ruby.

Responsibilities

  • Own SRE solutions end-to-end, from design and implementation to operation and continuous improvement.
  • Use Infrastructure-as-Code and configuration management to standardize and automate provisioning.
  • Deliver solutions in a globally distributed, multi-cloud hybrid environment.
  • Design for failure with redundancy, failure domains, and strict change control.
  • Ensure the highest level of uptime and Quality of Service for internal customers.
  • Conduct capacity management and planning to meet ongoing operational needs.
  • Detect performance issues and recommend solutions to maintain service quality.
  • Collaborate with various teams to ensure seamless project completion.
  • Participate in on-call, incident reviews, and assist in root cause identification.

Benefits

  • Employees at NVIDIA are often offered comprehensive, day-one benefits—including medical, dental, and vision coverage with HSA support, life and disability insurance, an Employee Assistance Program, and a 401(k) with auto-enrollment. Many roles also have generous time off and holidays, donation matching (up to $10,000), and a wide menu of extras like FSAs, commuter benefits, legal and identity-theft protection, pet insurance, and wellness discounts. Optional programs can include student-loan and home-purchase support, plus family care resources and expert medical services.

Is this posting expired or inaccurate?