JobsSite Reliability Engineer - Hardware Infrastructure
Site Reliability Engineer - Hardware Infrastructure
NVIDIASite Reliability Engineer - Hardware Infrastructure
NVIDIALocation
Santa Clara, CA
Type
Full-time
Posted
5/10/2026
Compensation
$184,000 - $356,500 per year
Undergraduate with 5+ Years of Experience
Approval 99.2%·Filings 1,781·New hires 873·
👑 Elite Sponsor
·FY 2025Job description
The Site Reliability Engineering team at NVIDIA offers an opportunity to define, develop, and support large-scale production systems with a focus on efficiency and availability. This role combines software and systems engineering to ensure reliable service operation. As an SRE, you will work in a collaborative environment that encourages creativity and empowers developers. Your contributions will help maintain system functionality while implementing significant updates.
Requirements
- Degree in Computer Science or a related technical field involving coding, or equivalent experience.
- 8+ years of experience in SRE, DevOps, or Production Engineering.
- Strong understanding of SRE principles, including incident management, error budgets, SLOs, and SLAs.
- Experience crafting and deploying systems that are fault-tolerant, performant, and supportable.
- Background with infrastructure automation.
- Experience running critical services in production.
- Experience in one or more of the following: Python, Go, Perl, or Ruby.
- Hands-on experience with observability platforms such as Prometheus or Grafana.
- Strong communication skills with the ability to convey technical concepts effectively to diverse audiences.
- Flexibility and adaptability working in a fast-paced environment with evolving requirements.
Responsibilities
- Develop and support guidelines for incident management, planned maintenance, and blameless postmortems.
- Assist teams in responding to high severity incidents, driving root cause analysis, crafting high-quality postmortems, and developing post-incident corrective actions.
- Define reliability and supportability metrics, Service Level Objectives, and error budgets.
- Develop and drive the adoption of actionable, customer-centric monitoring and alerting.
- Apply automation and Generative AI/Agentic solutions to minimize manual and tedious activities and boost customer support.
- Guide teams on establishing sustainable on-call and operational standards.
Benefits
- Employees at NVIDIA are often offered comprehensive, day-one benefits—including medical, dental, and vision coverage with HSA support, life and disability insurance, an Employee Assistance Program, and a 401(k) with auto-enrollment. Many roles also have generous time off and holidays, donation matching (up to $10,000), and a wide menu of extras like FSAs, commuter benefits, legal and identity-theft protection, pet insurance, and wellness discounts. Optional programs can include student-loan and home-purchase support, plus family care resources and expert medical services.
Is this posting expired or inaccurate?
