JobsSite Reliability Engineer (Edge Services), Infrastructure Services
Apple logo

Site Reliability Engineer (Edge Services), Infrastructure Services

Apple

Location

Austin, TX, Denver, CO, Elk Grove, CA, Sunnyvale, CA

Type

Full-time

Posted

5/19/2026

Compensation

Not listed

Undergraduate with 2+ Years of Experience
Approval 98.9%·Filings 5,543·New hires 2,691·
👑 Elite Sponsor
·FY 2025

Job description

We are looking for a proactive Site Reliability Engineer to enhance our production ecosystems by developing a sophisticated reliability framework. This role involves ensuring our services are resilient, scalable, and observable while treating operations as a software problem. The SRE team member will focus on building self-healing systems and reducing toil through automation. Additionally, the engineer will collaborate with development teams to integrate reliability into the CI/CD pipeline.

Requirements

  • Strong understanding of Linux internals and deep networking expertise, including HTTP/2, HTTP/3, and HTTPS/TLS.
  • Proven ability to automate repetitive tasks and complex workflows using Python or Go.
  • Experience configuring and managing modern monitoring suites such as Prometheus, Grafana, and ClickHouse.
  • Knowledge of Data Structures and Algorithms to write efficient code and troubleshoot system bottlenecks.
  • Practical knowledge of SLIs, SLOs, Error Budgets, Release Management, and Incident Management.
  • Experience managing cloud environments like AWS, GCP, or Azure using Terraform, Ansible, or Pulumi.
  • Hands-on experience scaling and securing containerized workloads via Kubernetes.
  • Ability to lead blameless post-mortems to improve system resilience.
  • Proactive engineering mindset focused on designing systems to prevent failures.
  • Practical fluency in applying Generative AI tools within SRE and software engineering workflows.

Responsibilities

  • Drive the vision for service visibility and build a data-driven reliability framework.
  • Design and implement a next-generation observability and alerting strategy.
  • Build self-healing systems and reduce operational toil through automation.
  • Partner with development teams to integrate reliability into the CI/CD pipeline.
  • Identify and mitigate performance bottlenecks proactively.
  • Consult with product teams on service design for long-term maintainability.
  • Lead blameless post-mortems to enhance system robustness.
  • Apply Generative AI tools to improve observability and debugging workflows.

Benefits

  • Employees at Apple are often offered comprehensive benefits that support physical and mental well-being—flexible medical plans, confidential counseling, onsite wellness centers at major campuses, and resources for fitness and daily life. Families typically receive fertility support, paid parental leave with gradual return, caregiving leave, and dependent-care guidance, while financial perks commonly include stock grants (with purchase discounts), 401(k) matching, and income-protection coverage. Employees also see robust time off, Apple University learning and tuition reimbursement, donation matching and paid volunteer hours, and deep product and partner discounts.

Is this posting expired or inaccurate?