JobsPost-Training Platform Infrastructure Engineer
Location
San Jose, CA
Type
Full-time
Posted
6/17/2026
Compensation
USD $178,500.00/Yr. – USD $255,000.00/Yr.
Undergraduate with 2+ Years of Experience
Approval 98.6%·Filings 728·New hires 184·
✓ Established Sponsor
·FY 2025Job description
The role at AMD is for a systems-minded engineer focused on large-scale model inference, distributed systems, and performance optimization. The engineer will work on post-training and inference infrastructure, emphasizing P/D disaggregation and KV cache management. The position requires a deep understanding of modern ML infrastructure and the ability to translate research insights into production-grade features. Collaboration with research and applied ML teams is essential to enhance system capabilities and performance.
Requirements
- Strong background in systems engineering, distributed systems, or ML infrastructure.
- Hands-on experience with GPU-accelerated workloads and memory-constrained systems.
- Proficiency in Python and C++ or similar systems languages.
- Ability to read, understand, and modify complex open-source codebases.
- Direct experience with LLM inference frameworks or serving stacks.
Responsibilities
- Research and understand modern LLM inference frameworks.
- Analyze and compare inference execution paths to identify performance bottlenecks.
- Develop and implement infrastructure-level features to improve inference latency and memory efficiency.
- Collaborate with research and applied ML teams to translate model-level requirements into infrastructure capabilities.
- Document findings, architectural insights, and best practices to guide future system design.
Benefits
- AMD provides a competitive 'Total Rewards' package that focuses on financial growth, health, and work-life balance.
Is this posting expired or inaccurate?
