JobsSenior Deep Learning Architect, LLM Inference
NVIDIA logo

Senior Deep Learning Architect, LLM Inference

NVIDIA

Location

Santa Clara, CA

Type

Full-time

Posted

8/18/2026

Compensation

$184,000 - $356,500 per year

PhD with 5+ Years of Experience
Approval 99.2%·Filings 1,781·New hires 873·
👑 Elite Sponsor
·FY 2025

Job description

The Senior Deep Learning Architect for LLM Inference at NVIDIA will focus on optimizing inference server performance for Large Language Models. This role is part of the Inference Benchmarking team, which is dedicated to maintaining NVIDIA's leadership in generative AI. The architect will collaborate with various teams, including performance marketing and AI startups, to develop benchmarking methodologies and tools. A strong background in deep learning and software development is essential for driving advancements in this rapidly evolving field.

Requirements

  • Master's or PhD degree in Computer Science, Computer Engineering, related fields, or equivalent experience.
  • 6+ years of relevant software development experience.
  • Detailed knowledge of deep learning inference serving, PyTorch programming, profiling, and compiler optimizations.
  • Experience developing client server LLM applications with OpenAI API or MCP and identifying performance bottlenecks.
  • Solid understanding of CPU and GPU microarchitecture and performance characteristics.
  • Experience with complex software projects like frameworks, compilers, or operating systems.
  • Demonstrated proficiency with the latest AI coding agents like Claude Code, Codex, and Cursor.
  • Excellent written and verbal communication skills and the ability to work independently and collaboratively in a fast-paced environment.

Responsibilities

  • Conduct workload characterization of the latest LLMs and inference servers to ensure NVIDIA maintains its leadership position.
  • Collaborate with the performance marketing team to create engaging content highlighting NVIDIA's inference achievements.
  • Work with engineers from AI startup companies to establish standard benchmarking methodologies.
  • Develop a constantly evolving inference performance data results website.
  • Invent end-to-end profiling and analysis tools to keep up with the rapid pace of Generative AI.
  • Contribute to deep learning software projects to drive advancements in the field.
  • Verify that new GPU product launches produce industry-leading performance.
  • Collaborate across the company to guide the direction of inference serving.

Benefits

  • Employees at NVIDIA are often offered comprehensive, day-one benefits—including medical, dental, and vision coverage with HSA support, life and disability insurance, an Employee Assistance Program, and a 401(k) with auto-enrollment. Many roles also have generous time off and holidays, donation matching (up to $10,000), and a wide menu of extras like FSAs, commuter benefits, legal and identity-theft protection, pet insurance, and wellness discounts. Optional programs can include student-loan and home-purchase support, plus family care resources and expert medical services.

Is this posting expired or inaccurate?