Job description
NVIDIA is seeking an NCX Senior Engineer to join the DSX team, focusing on NVIDIA Cloud Partner infrastructure operations. This role emphasizes enhancing operational capabilities for large-scale NVIDIA accelerated infrastructure in production. The engineer will guide partners through advanced Day 2 operations, ensuring ongoing infrastructure health and performance validation. This position requires a hands-on approach at the intersection of accelerated computing, cloud infrastructure, and production operations.
Requirements
- BS, MS, or Ph.D. in Computer Science, Computer/Electrical Engineering, or a related technical field, or equivalent experience.
- 8+ years of experience in infrastructure engineering, Site Reliability Engineering, DevOps, cloud platform engineering, systems engineering, or similar roles supporting large-scale production environments.
- Strong experience operating Linux-based distributed systems and cloud infrastructure in production.
- Deep understanding of Kubernetes, containers, cluster scheduling, and the operational lifecycle of large multi-node environments.
- Strong understanding of production observability, including metrics, logging, alerting, dashboards, health checks, and operations guided by service level agreements.
- Experience crafting automation for infrastructure lifecycle management, failure detection, remediation, upgrades, and configuration management.
- Strong networking fundamentals and experience troubleshooting complex distributed systems across compute, network, and storage layers.
- Programming and automation experience using Python, Go, shell scripting, or similar languages.
Responsibilities
- Lead NCP Day 2 operational readiness efforts by collaborating with NVIDIA Cloud Partners to set up necessary systems and procedures.
- Develop and implement methods to continuously validate GPU, CPU, storage, and network health across large-scale AI clusters.
- Help NCPs implement comprehensive telemetry, monitoring, alerting, dashboards, and operational signals across various workloads.
- Build workflows to detect, isolate, drain, repair, validate, and return unhealthy infrastructure to service.
- Implement scalable strategies for managing sizable GPU fleets, including lifecycle administration and configuration management.
- Translate NVIDIA NCP requirements and reference architectures into production operating practices and measurable operational standards.
- Develop health signals, SLOs, important metrics, acceptance criteria, and ongoing validation mechanisms for infrastructure reliability.
- Create tooling, automation, implementation guides, runbooks, operational playbooks, and reference implementations for consistent application across NCP environments.
Benefits
- Employees at NVIDIA are often offered comprehensive, day-one benefits—including medical, dental, and vision coverage with HSA support, life and disability insurance, an Employee Assistance Program, and a 401(k) with auto-enrollment. Many roles also have generous time off and holidays, donation matching (up to $10,000), and a wide menu of extras like FSAs, commuter benefits, legal and identity-theft protection, pet insurance, and wellness discounts. Optional programs can include student-loan and home-purchase support, plus family care resources and expert medical services.
H-1B filing history
Public USCIS petition and DOL LCA counts · latest USCIS FY2026, LCA FY2026
Filing entity: Nvidia Corporation
As of Aug 23, 2026
Initial approvals
355
FY2026
Approval rate
99.2%
FY2026
LCA certified
7,667
FY2026
Entry-level share
11.9%
FY2026
Initial approvals YoY
-37%
Trend
LCA certified YoY
-83%
Trend
Initial approvals by fiscal year
Approval rate by fiscal year
Continuing vs initial approvals
LCA certified positions by quarter
LCA certified positions by fiscal year
Based on public USCIS and DOL filings; not a sponsorship guarantee.
Is this posting expired or inaccurate?
