JobsSenior Software Development Engineer in Test â LLM Evaluation & Automation, T3E
Senior Software Development Engineer in Test â LLM Evaluation & Automation, T3E
AppleSenior Software Development Engineer in Test â LLM Evaluation & Automation, T3E
AppleLocation
Cupertino, CA
Type
Full-time
Posted
8/4/2026
Compensation
Not listed
Undergraduate with 5+ Years of Experience
Approval 98.9%·Filings 5,543·New hires 2,691·
👑 Elite Sponsor
·FY 2025Job description
The Apple Intelligence Platform Experience Validation team is seeking a Senior SDET to lead the design and implementation of automated model evaluation. This hands-on role focuses on ensuring high-quality Apple Intelligence features by building and maintaining evaluation automation across various generative features. The candidate will partner with modeling, framework, and infrastructure teams to create reliable evaluation processes that catch model regressions. The position requires strong technical skills and the ability to communicate effectively with diverse teams.
Requirements
- BS in Computer Science, Mathematics, or a related field or equivalent practical experience.
- Three years of relevant industry experience in test automation, software development, or related areas.
- Strong practical knowledge of Python, including data-pipeline fluency with JSON/YAML and REST APIs.
- Hands-on experience with LLM-as-a-judge evaluation and rubric design.
- Strong software engineering fundamentals with the ability to define atomic, composable components.
- Strong debugging and triage skills to separate genuine regressions from infrastructure noise.
- Strong knowledge of the software development lifecycle, testing methodologies, and QA processes.
- Excellent written and verbal communication skills.
- Ability to lead work across varying priorities and partner multi-functionally.
- Experience building on-device tooling and device/model evaluation infrastructure.
- Experience integrating with CI/CD and job orchestration systems.
- Familiarity with generative model behavior, including image generation and NLP.
- Experience curating and reasoning about large datasets.
- Awareness of dataset bias and fairness considerations in evaluation.
- Experience with database/query tooling and dashboards for reporting quality trends.
- Experience with Xcode is a bonus.
Responsibilities
- Lead the design and implementation of automated model evaluation processes.
- Build and maintain model level, component, or end-to-end evaluation coverage.
- Leverage LLM judge scoring output quality in automation.
- Validate model output and associated classification metadata for image generation.
- Evaluate generated artifacts and responses for natural-language generation.
- Replace exact-match checks with an LLM-as-judge stage integrated into the pipeline.
- Assess the sensibility and quality of model-generated content.
- Decide when component-level checks are sufficient versus full end-to-end user flow requirements.
- Build tooling for both component-level checks and end-to-end evaluations.
Benefits
- Employees at Apple are often offered comprehensive benefits that support physical and mental well-being—flexible medical plans, confidential counseling, onsite wellness centers at major campuses, and resources for fitness and daily life. Families typically receive fertility support, paid parental leave with gradual return, caregiving leave, and dependent-care guidance, while financial perks commonly include stock grants (with purchase discounts), 401(k) matching, and income-protection coverage. Employees also see robust time off, Apple University learning and tuition reimbursement, donation matching and paid volunteer hours, and deep product and partner discounts.
Is this posting expired or inaccurate?
