JobsSite Reliability Engineer - Hardware Infrastructure
NVIDIA logo

Site Reliability Engineer - Hardware Infrastructure

NVIDIA

Location

Santa Clara, CA

Type

Full-time

Posted

5/10/2026

Compensation

$184,000 - $356,500 per year

Undergraduate with 5+ Years of Experience
Approval 99.2%·Filings 1,781·New hires 873·
👑 Elite Sponsor
·FY 2025

Job description

The Site Reliability Engineering team at NVIDIA offers an opportunity to define, develop, and support large-scale production systems with a focus on efficiency and availability. This role combines software and systems engineering to ensure reliable service operation. As an SRE, you will work in a collaborative environment that encourages creativity and empowers developers. Your contributions will help maintain system functionality while implementing significant updates.

Requirements

  • Degree in Computer Science or a related technical field involving coding, or equivalent experience.
  • 8+ years of experience in SRE, DevOps, or Production Engineering.
  • Strong understanding of SRE principles, including incident management, error budgets, SLOs, and SLAs.
  • Experience crafting and deploying systems that are fault-tolerant, performant, and supportable.
  • Background with infrastructure automation.
  • Experience running critical services in production.
  • Experience in one or more of the following: Python, Go, Perl, or Ruby.
  • Hands-on experience with observability platforms such as Prometheus or Grafana.
  • Strong communication skills with the ability to convey technical concepts effectively to diverse audiences.
  • Flexibility and adaptability working in a fast-paced environment with evolving requirements.

Responsibilities

  • Develop and support guidelines for incident management, planned maintenance, and blameless postmortems.
  • Assist teams in responding to high severity incidents, driving root cause analysis, crafting high-quality postmortems, and developing post-incident corrective actions.
  • Define reliability and supportability metrics, Service Level Objectives, and error budgets.
  • Develop and drive the adoption of actionable, customer-centric monitoring and alerting.
  • Apply automation and Generative AI/Agentic solutions to minimize manual and tedious activities and boost customer support.
  • Guide teams on establishing sustainable on-call and operational standards.

Benefits

  • Employees at NVIDIA are often offered comprehensive, day-one benefits—including medical, dental, and vision coverage with HSA support, life and disability insurance, an Employee Assistance Program, and a 401(k) with auto-enrollment. Many roles also have generous time off and holidays, donation matching (up to $10,000), and a wide menu of extras like FSAs, commuter benefits, legal and identity-theft protection, pet insurance, and wellness discounts. Optional programs can include student-loan and home-purchase support, plus family care resources and expert medical services.

Is this posting expired or inaccurate?