JobsPrincipal ML Engineer - Large Scale Training Performance Optimization
Principal ML Engineer - Large Scale Training Performance Optimization
AMDPrincipal ML Engineer - Large Scale Training Performance Optimization
AMDLocation
San Jose, CA
Type
Full-time
Posted
5/5/2026
Compensation
USD $210,000.00/Yr. – USD $300,000.00/Yr.
PhD with 5+ Years of Experience
Approval 98.6%·Filings 728·New hires 184·
✓ Established Sponsor
·FY 2025Job description
The Principal Machine Learning Engineer will join AMD's Models and Applications team, focusing on the distributed training of large models on multiple GPUs. This role emphasizes improving training efficiency and innovating within the realm of generative AI. The ideal candidate will contribute to a collaborative environment that values diverse perspectives and bold ideas. This position offers an opportunity to influence the direction of AMD's AI platform while working with a world-class team.
Requirements
- Experience with distributed training pipelines.
- Knowledge of distributed training algorithms such as Data Parallel, Tensor Parallel, Pipeline Parallel, and Expert Parallel ZeRO.
- Familiarity with training large models at scale.
- Excellent programming skills in Python or C++, including debugging and performance analysis.
Responsibilities
- Train large models to convergence on AMD GPUs at scale.
- Improve the end-to-end training pipeline performance.
- Optimize the distributed training pipeline and algorithm to scale out.
- Contribute changes to open source projects.
- Stay up-to-date with the latest training algorithms.
- Influence the direction of AMD's AI platform.
- Collaborate across teams with various groups and stakeholders.
Benefits
- AMD provides a competitive 'Total Rewards' package that focuses on financial growth, health, and work-life balance.
Is this posting expired or inaccurate?
