Location
remote, Austin, TX
Type
Full-time
Posted
9/13/2026
Compensation
$108,000 - $207,000 per year
Job description
The Technical Support Engineer at NVIDIA will focus on supporting Slurm, a critical workload manager used in AI and high-performance computing environments. This role involves diagnosing complex issues and providing solutions to customers running production clusters. The engineer will work closely with a dedicated team to ensure reliable and efficient operations. A strong background in Slurm and Linux systems is essential for success in this position.
Requirements
- BS degree in Computer Science, Engineering, or a related field, or equivalent experience.
- 5+ years of hands-on experience administering and supporting Slurm in production HPC or AI environments.
- Expert-level understanding of Slurm architecture, daemons, configuration, scheduling behavior, accounting, resource management, and failure modes.
- In-depth Linux system-administration experience, including systemd, cgroups, authentication, networking, and database-backed services.
- Strong analytical and research skills with proficiency in distinguishing Slurm defects from configuration, integration, infrastructure, and workload problems.
- Excellent written and verbal communication skills.
Responsibilities
- Own Slurm support cases from initial investigation through resolution for customers running production AI and HPC clusters.
- Diagnose complex problems involving slurmctld, slurmd, slurmdbd, job scheduling, node management, resource allocation, accounting, authentication, and high availability.
- Solve Slurm configuration and policy features, including partitions, reservations, priorities, fair-share, quality of service, backfill, preemption, GRES/TRES, cgroups, and job constraints.
- Investigate performance, reliability, and scalability issues using logs, diagnostic data, configuration analysis, reproductions, and source-level debugging.
- Isolate problems across Slurm and its surrounding dependencies, including Linux, MUNGE, databases, networking, parallel storage, containers, GPUs, and cluster-management systems.
- Advise customers on Slurm configuration, upgrades, operational practices, managing system resources, and safe recovery from production incidents.
- Collaborate with engineering teams by producing clear technical descriptions, reproducible test cases, and well-supported defect reports.
- Develop guides, knowledge-base articles, diagnostic tools, and internal training that strengthen Slurm expertise across the support organization.
Benefits
- Employees at NVIDIA are often offered comprehensive, day-one benefits—including medical, dental, and vision coverage with HSA support, life and disability insurance, an Employee Assistance Program, and a 401(k) with auto-enrollment. Many roles also have generous time off and holidays, donation matching (up to $10,000), and a wide menu of extras like FSAs, commuter benefits, legal and identity-theft protection, pet insurance, and wellness discounts. Optional programs can include student-loan and home-purchase support, plus family care resources and expert medical services.
H-1B filing history
Public USCIS petition and DOL LCA counts · latest USCIS FY2026, LCA FY2026
Filing entity: Nvidia Corporation
As of Aug 23, 2026
Initial approvals
355
FY2026
Approval rate
99.2%
FY2026
LCA certified
7,667
FY2026
Entry-level share
11.9%
FY2026
Initial approvals YoY
-37%
Trend
LCA certified YoY
-83%
Trend
Initial approvals by fiscal year
Approval rate by fiscal year
Continuing vs initial approvals
LCA certified positions by quarter
LCA certified positions by fiscal year
Based on public USCIS and DOL filings; not a sponsorship guarantee.
Is this posting expired or inaccurate?
