JobsStaff Software Engineer - GenAI inference
Databricks logo

Staff Software Engineer - GenAI inference

Databricks

Location

San Francisco, CA

Type

Full-time

Posted

8/25/2026

Compensation

Not listed

Undergraduate with 5+ Years of Experience
H-1B FY202699.7% approval+19% YoY
💎 Strong sponsor

Job description

As a staff software engineer for GenAI inference at Databricks, you will lead the architecture and optimization of the inference engine for the Databricks Foundation Model API. This role involves bridging research advances with production demands to ensure high throughput and low latency. You will work on the full GenAI inference stack, collaborating closely with researchers and cross-functional teams. Your contributions will be critical in driving performance and reliability in machine learning inference systems.

Requirements

  • BS/MS/PhD in Computer Science or a related field
  • Strong software engineering background with 6+ years of experience in performance-critical systems
  • Proven track record of owning complex system components and driving architectural decisions end-to-end
  • Deep understanding of ML inference internals, including attention, MLPs, and quantization
  • Hands-on experience with CUDA, GPU programming, and key libraries such as cuBLAS and cuDNN
  • Strong background in distributed systems design, including RPC frameworks and memory partitioning
  • Demonstrated ability to uncover and solve performance bottlenecks across layers
  • Experience building instrumentation, tracing, and profiling tools for ML models
  • Excellent communication and leadership skills with a proactive mindset

Responsibilities

  • Own and drive the architecture, design, and implementation of the inference engine.
  • Collaborate with researchers to integrate new model architectures or features into the engine.
  • Lead the end-to-end optimization for latency, throughput, and memory efficiency.
  • Define and guide standards for instrumentation, profiling, and tracing tooling.
  • Architect scalable routing, batching, scheduling, and memory management mechanisms.
  • Ensure reliability and fault tolerance in the inference pipelines.
  • Collaborate on integrating with federated and distributed inference infrastructure.
  • Drive cross-team collaboration with platform engineers and security teams.
  • Represent the team externally through benchmarks and open-source contributions.

H-1B filing history

Public USCIS petition and DOL LCA counts · latest USCIS FY2026, LCA FY2026

Filing entity: Databricks Inc

As of Aug 23, 2026

Initial approvals

83

FY2026

Approval rate

99.7%

FY2026

LCA certified

174

FY2026

Entry-level share

4.0%

FY2026

Initial approvals YoY

+19%

Trend

LCA certified YoY

-54%

Trend

Initial approvals by fiscal year

Approval rate by fiscal year

Continuing vs initial approvals

LCA certified positions by quarter

LCA certified positions by fiscal year

Based on public USCIS and DOL filings; not a sponsorship guarantee.

Is this posting expired or inaccurate?