JobsMember of Technical Staff, Site Reliability Engineer (HPC) - MAI SuperIntelligence Team
Microsoft logo

Member of Technical Staff, Site Reliability Engineer (HPC) - MAI SuperIntelligence Team

Microsoft

Location

Mountain View, CA

Type

Full-time

Posted

5/5/2026

Compensation

$119,800 - $304,200 per year

Master's with 2+ Years of Experience
Undergraduate with 4+ Years of Experience
H-1B FY202696.0% approval-64% YoY
👑 Elite sponsor

Job description

The role of Site Reliability Engineer (SRE) focuses on maintaining the reliability and efficiency of Microsoft's High Performance Computing (HPC) infrastructure. This position is part of the Superintelligence Team, which aims to advance AI technologies while ensuring they align with human values. The SRE will blend software and systems engineering to support large-scale distributed AI systems. The team values collaboration, innovation, and a growth mindset to empower users and organizations globally.

Requirements

  • Master's Degree in Computer Science, Information Technology, or related field AND 2+ years technical experience in Site Reliability Engineering, DevOps, or Infrastructure Engineering.
  • OR Bachelor's Degree in Computer Science, Information Technology, or related field AND 4+ years technical experience in Site Reliability Engineering, DevOps, or Infrastructure Engineering.
  • Strong proficiency in Kubernetes, Docker, and container orchestration.
  • Knowledge of CI/CD pipelines for Inference and ML model deployment.
  • Hands-on experience with public cloud platforms like Azure/AWS/GCP and infrastructure-as-code.
  • Expertise in monitoring and observability tools such as Grafana, Datadog, and OpenTelemetry.
  • Strong programming/scripting skills in Python, Go, or Bash.
  • Solid knowledge of distributed systems, networking, and storage.
  • Experience running large-scale GPU clusters for ML/AI workloads.
  • Familiarity with ML training/inference pipelines.
  • Experience with high-performance computing (HPC) and workload schedulers.
  • Background in capacity planning and cost optimization for GPU-heavy environments.

Responsibilities

  • Ensure uptime, resiliency, and fault tolerance of HPC clusters powering MAI model training and inference.
  • Design and maintain monitoring, alerting, and logging systems for real-time visibility into HPC systems.
  • Build automation for deployments, incident response, scaling, and failover in CPU+GPU environments.
  • Lead on-call rotations, troubleshoot production issues, and conduct blameless postmortems.
  • Ensure data privacy, compliance, and secure operations across model training and serving environments.
  • Partner with ML engineers and platform teams to improve developer experience and accelerate research-to-production workflows.

Benefits

  • Employees at Microsoft are often offered comprehensive, “world-class” benefits—including health and mental-wellness programs, competitive pay with bonuses and stock awards, and retirement/savings options. Time-off and flexibility are common, with generous vacation and holidays, parental and caregiver leave, and flexible work schedules, alongside learning support, employee resource groups, product discounts, and matching-gifts/volunteering programs. Specific benefits can vary by region.

H-1B filing history

Public USCIS petition and DOL LCA counts · latest USCIS FY2026, LCA FY2026

Filing entity: Microsoft Corporation

As of Aug 23, 2026

Initial approvals

503

FY2026

Approval rate

96.0%

FY2026

LCA certified

3,062

FY2026

Entry-level share

5.7%

FY2026

Initial approvals YoY

-64%

Trend

LCA certified YoY

-83%

Trend

Initial approvals by fiscal year

Approval rate by fiscal year

Continuing vs initial approvals

LCA certified positions by quarter

LCA certified positions by fiscal year

Based on public USCIS and DOL filings; not a sponsorship guarantee.

Is this posting expired or inaccurate?