Senior Software Engineer - AI Inference Performance
NVIDIASenior Software Engineer - AI Inference Performance
NVIDIALocation
Santa Clara, CA
Type
Full-time
Posted
9/29/2026
Compensation
$184,000 - $356,500 per year
Job description
NVIDIA is seeking a Senior Software Engineer – AI Inference Performance to enhance LLM and VLM inference on GPU-accelerated systems. This hands-on role involves optimizing performance metrics such as latency and throughput while collaborating with various teams. The engineer will work on performance-critical kernels and contribute to open-source inference engines. The position requires a strong understanding of GPU architecture and performance optimization techniques.
Requirements
- More than 6 years of experience in full-stack LLM/VLM inference performance involving models, serving, distributed runtimes, kernels, and hardware.
- Strong programming skills in Python, Rust, and/or C++, plus hands-on experience with CUDA or another GPU programming environment.
- Demonstrated expertise in speed-of-light analysis, roofline models, microbenchmarks, and tools including NVIDIA Nsight Systems and Nsight Compute.
- Deep understanding of GPU architecture, including Tensor Cores, memory hierarchy, caches, occupancy, synchronization, and numerical formats across hardware generations.
- Practical experience optimizing inference servers and model execution.
- Understanding of distributed systems and networking for accelerated computing.
- BS or MS in Computer Science, Computer Engineering, or a related field, or equivalent experience.
Responsibilities
- Lead end-to-end analysis of LLM/VLM inference processes.
- Define representative prefill and decode workloads.
- Optimize time to first token, inter-token latency, P99 end-to-end latency, processing efficiency, and key-value cache capacity.
- Build speed-of-light and roofline models to quantify performance headroom.
- Profile workloads using NVIDIA Nsight Systems, Nsight Compute, PyTorch Profiler, and custom instrumentation.
- Tune serving hyperparameters and techniques such as batching, KV-cache management, quantization, speculative decoding, CUDA Graphs, and model parallelism.
- Build and optimize performance-critical kernels, including attention, matrix multiplication, mixture-of-experts routing, quantization, and data movement.
- Establish repeatable benchmarks, canonical run records, and performance regression gates.
Benefits
- Employees at NVIDIA are often offered comprehensive, day-one benefits—including medical, dental, and vision coverage with HSA support, life and disability insurance, an Employee Assistance Program, and a 401(k) with auto-enrollment. Many roles also have generous time off and holidays, donation matching (up to $10,000), and a wide menu of extras like FSAs, commuter benefits, legal and identity-theft protection, pet insurance, and wellness discounts. Optional programs can include student-loan and home-purchase support, plus family care resources and expert medical services.
H-1B filing history
Public USCIS petition and DOL LCA counts · latest USCIS FY2026, LCA FY2026
Filing entity: Nvidia Corporation
As of Aug 23, 2026
Initial approvals
355
FY2026
Approval rate
99.2%
FY2026
LCA certified
7,667
FY2026
Entry-level share
11.9%
FY2026
Initial approvals YoY
-37%
Trend
LCA certified YoY
-83%
Trend
Initial approvals by fiscal year
Approval rate by fiscal year
Continuing vs initial approvals
LCA certified positions by quarter
LCA certified positions by fiscal year
Based on public USCIS and DOL filings; not a sponsorship guarantee.
Is this posting expired or inaccurate?
