JobsPrincipal Site Reliability Engineering Manager- CTJ- Secret (Cleared Environments)
Principal Site Reliability Engineering Manager- CTJ- Secret (Cleared Environments)
MicrosoftPrincipal Site Reliability Engineering Manager- CTJ- Secret (Cleared Environments)
MicrosoftLocation
Redmond, WA
Type
Full-time
Posted
5/5/2026
Compensation
$139,900 - $304,200 per year
PhD with 5+ Years of Experience
Master's with 5+ Years of Experience
Undergraduate with 5+ Years of Experience
Approval 98.4%·Filings 6,363·New hires 3,142·
👑 Elite Sponsor
·FY 2025Job description
The Principal Site Reliability Engineering Manager will lead a team responsible for building and operating Substrate services in highly regulated environments. This role emphasizes the importance of strong software engineering fundamentals to ensure operational excellence and reliability. The manager will develop senior engineers and influence engineering strategy while ensuring compliance and security are integrated into service design. The position requires hands-on leadership during incidents and a focus on continuous improvement of reliability and customer experience.
Requirements
- Doctorate Degree in Computer Science, Information Technology, or related field AND 2+ years technical experience in software engineering, network engineering, or systems administration OR Master's Degree in Computer Science, Information Technology, or related field AND 3+ years technical experience OR Bachelor's Degree in Computer Science, Information Technology, or related field AND 5+ years technical experience OR equivalent experience.
- Candidates must be able to meet Microsoft, customer and/or government security screening requirements.
- This role requires access to Microsoft Government cloud environments, including GCC Moderate, GCC High, and Department of Defense environments.
Responsibilities
- Lead and develop a team of Site Reliability Engineer ICs, providing clear expectations, regular coaching, and career guidance.
- Own the operational health and reliability posture of Substrate services running in regulated environments.
- Drive change and influence across the organization by establishing and driving SLOs, SLIs, and operational metrics.
- Lead effective incident management and post-incident reviews, emphasizing systemic fixes and long-term resilience.
- Serve as an actively engaged on-call engineer and participate in an on-call rotation.
- Own reliability, resilience, and disaster recovery, including driving and coordinating DR and game day exercises.
- Drive engineering-led operational excellence at scale, leveraging strong software engineering practices and automation.
- Partner with engineering and product teams to embed reliability, security, and compliance considerations early in service design.
- Influence technical and operational strategy beyond your immediate team, particularly for cross-cutting reliability and compliance challenges.
- Represent your team’s work clearly to leadership and partners, articulating risks, trade-offs, and progress.
Benefits
- Employees at Microsoft are often offered comprehensive, “world-class” benefits—including health and mental-wellness programs, competitive pay with bonuses and stock awards, and retirement/savings options. Time-off and flexibility are common, with generous vacation and holidays, parental and caregiver leave, and flexible work schedules, alongside learning support, employee resource groups, product discounts, and matching-gifts/volunteering programs. Specific benefits can vary by region.
Is this posting expired or inaccurate?
