JobsPrincipal Site Reliability Engineer
NVIDIA logo

Principal Site Reliability Engineer

NVIDIA

Location

Santa Clara, CA

Type

Full-time

Posted

9/9/2026

Compensation

$248,000 - $396,750 per year

Undergraduate with 5+ Years of Experience
H-1B FY202699.2% approval-37% YoY
👑 Elite sponsor

Job description

The Principal Site Reliability Engineer (SRE) at NVIDIA will lead the technical direction of the AI Platform Runtime and oversee reliability engineering initiatives across various organizations. This role focuses on solving complex problems at an enterprise scale while establishing architectural standards for critical AI-powered services. The SRE team emphasizes collaboration, intellectual curiosity, and continuous improvement in operational efficiency and system reliability. The position requires a blend of software and systems engineering expertise to enhance platform operations and developer productivity.

Requirements

  • 15+ years of experience in Site Reliability Engineering, Platform Engineering, Distributed Systems, Cloud Architecture, or related infrastructure engineering roles.
  • BS or MS degree in Computer Science or a related technical field involving significant software development, or equivalent experience.
  • Demonstrated experience setting technical strategy and leading large-scale engineering initiatives across multiple teams or organizations.
  • Deep expertise in distributed systems architecture, networking, Linux, Kubernetes, and public cloud platforms such as AWS, Azure, or GCP.
  • Strong proficiency in one or more programming languages such as Python, Go, TypeScript, JavaScript, or Java.
  • Extensive experience with infrastructure-as-code and platform automation technologies such as Terraform, Crossplane, AWS CDK, or AWS CloudFormation.
  • Deep understanding of observability at scale, including OpenTelemetry and modern metrics, logging, tracing, profiling, and analytics platforms.
  • Expertise in reliability engineering practices, including service-level objectives, error budgets, capacity planning, fault tolerance, disaster recovery, and incident management.
  • A track record of delivering measurable improvements in availability, performance, operational efficiency, or engineering productivity within complex environments.
  • Outstanding technical judgment, communication, and collaboration skills.

Responsibilities

  • Define and drive the long-term technical vision, architecture, and roadmap for reliability across NVIDIA’s AI Platform Runtime and related enterprise systems.
  • Architect highly available, resilient, secure, and scalable distributed platforms that power NVIDIA’s next-generation AI-driven products and services.
  • Lead the design and development of AI agents, AI skills, and intelligent automation that accelerate platform operations, incident response, troubleshooting, and remediation.
  • Establish platform-wide reliability standards, including service-level objectives, error budgets, capacity models, resilience patterns, and operational readiness requirements.
  • Identify systemic risks and lead cross-functional programs that improve availability, scalability, performance, security, and developer productivity.
  • Advance observability across complex distributed environments through OpenTelemetry, metrics, logs, traces, profiling, analytics, and automated anomaly detection.
  • Partner with senior leaders and engineers across Cloud, Platform, Security, Networking, and AI/ML organizations to make high-impact architectural and investment decisions.
  • Provide technical leadership during critical incidents and ensure that lessons from incidents translate into durable engineering and organizational improvements.
  • Develop reference architectures, shared platform capabilities, and automation frameworks that can be adopted across multiple engineering organizations.
  • Influence, mentor, and raise the technical bar for senior engineers and technical leaders throughout the company.

Benefits

  • Employees at NVIDIA are often offered comprehensive, day-one benefits—including medical, dental, and vision coverage with HSA support, life and disability insurance, an Employee Assistance Program, and a 401(k) with auto-enrollment. Many roles also have generous time off and holidays, donation matching (up to $10,000), and a wide menu of extras like FSAs, commuter benefits, legal and identity-theft protection, pet insurance, and wellness discounts. Optional programs can include student-loan and home-purchase support, plus family care resources and expert medical services.

H-1B filing history

Public USCIS petition and DOL LCA counts · latest USCIS FY2026, LCA FY2026

Filing entity: Nvidia Corporation

As of Aug 23, 2026

Initial approvals

355

FY2026

Approval rate

99.2%

FY2026

LCA certified

7,667

FY2026

Entry-level share

11.9%

FY2026

Initial approvals YoY

-37%

Trend

LCA certified YoY

-83%

Trend

Initial approvals by fiscal year

Approval rate by fiscal year

Continuing vs initial approvals

LCA certified positions by quarter

LCA certified positions by fiscal year

Based on public USCIS and DOL filings; not a sponsorship guarantee.

Is this posting expired or inaccurate?