JobsSite Reliability Engineer, Apple Data Platform - AI/ML Platform
Site Reliability Engineer, Apple Data Platform - AI/ML Platform
AppleSite Reliability Engineer, Apple Data Platform - AI/ML Platform
AppleLocation
Austin, TX
Type
Full-time
Posted
8/4/2026
Compensation
Not listed
Undergraduate with 2+ Years of Experience
Approval 98.9%·Filings 5,543·New hires 2,691·
👑 Elite Sponsor
·FY 2025Job description
The Apple Data Platform SRE team is responsible for maintaining a multi-cloud platform that supports internal engineers working on data and AI products. This role focuses on ensuring the reliability of services and infrastructure that power Apple's machine learning and AI initiatives. The ideal candidate will thrive on ownership and collaboration, driving their reliability roadmap while aligning with the team's broader goals. This position offers the opportunity to become an expert in cutting-edge technologies and contribute to the evolution of Apple's ML/AI infrastructure.
Requirements
- Bachelor's Degree in Computer Science, an engineering-related field, or equivalent related experience.
- 1-4 years in a Site Reliability Engineering, DevOps, or Infrastructure-focused role.
- Proficient in Python; working knowledge of Golang is a plus.
- Experience with Kubernetes and at least one major cloud provider (AWS or GCP).
- Exposure to operating or supporting ML pipelines, model-serving infrastructure, or LLM-based systems in production.
- Strong communication skills and composure under pressure during incidents.
- Solid grounding in SRE principles, with prior on-call or production-support experience.
- Hands-on experience operating or supporting Ray, LangGraph or similar agent orchestration frameworks, RAG architectures, embeddings platforms, or vector store platforms.
- Familiarity with MCP-based tooling and ML lifecycle/dataset management systems.
- Experience with S3 and cloud storage/networking fundamentals.
- Familiarity with observability tooling: Prometheus, Grafana, Splunk, PagerDuty.
- Working knowledge of CI/CD pipelines and deployment workflows.
- Deep understanding of one or more Big Data technologies (Spark, Flink, Airflow, Trino, Notebooks).
- A track record of automating manual operations through scripting or tooling.
- Intellectual curiosity and a drive to keep learning.
Responsibilities
- Operate and support the team's full portfolio, from big data pipelines to multi-cloud infrastructure.
- Ensure the reliability of machine learning and AI platform services.
- Collaborate with developers to maintain cutting-edge services and platforms.
- Drive the reliability roadmap for assigned services.
- Provide hands-on support to internal teams during incidents.
- Automate manual operations through scripting or tooling.
Benefits
- Employees at Apple are often offered comprehensive benefits that support physical and mental well-being—flexible medical plans, confidential counseling, onsite wellness centers at major campuses, and resources for fitness and daily life. Families typically receive fertility support, paid parental leave with gradual return, caregiving leave, and dependent-care guidance, while financial perks commonly include stock grants (with purchase discounts), 401(k) matching, and income-protection coverage. Employees also see robust time off, Apple University learning and tuition reimbursement, donation matching and paid volunteer hours, and deep product and partner discounts.
Is this posting expired or inaccurate?
