JobsAI/ML Platform Engineer
Job description
The AI / ML Platform Engineer role at AMD focuses on building a scalable and reliable platform for AI-for-engineering workflows. This position involves working closely with ML Systems Research Engineers, AI Research Scientists, and Applied AI Engineers to operationalize the Blueprint framework. The main emphasis is on developing infrastructure that supports large-scale agent execution, distributed training, and experiment tracking. Candidates should be strong systems engineers who can enhance reproducibility and performance across various workflows.
Requirements
- Strong programming skills in Python and one or more systems languages such as C++, Go, or Rust.
- Experience building ML platforms, AI infrastructure, or distributed systems.
- Strong understanding of job scheduling, distributed workloads, and production operations.
- Experience with Kubernetes, Ray, Slurm, or CI/CD processes.
- Good collaboration skills with AI researchers and applied engineers.
Responsibilities
- Build and operate the shared AI platform for agentic engineering workflows.
- Develop reliable infrastructure for distributed training and inference across GPU clusters.
- Build platform services for benchmark execution and reproducible evaluation.
- Maintain artifact systems for generated kernels and verification artifacts.
- Improve GPU cluster utilization and scheduling efficiency.
Benefits
- AMD provides a competitive 'Total Rewards' package that focuses on financial growth, health, and work-life balance.
Is this posting expired or inaccurate?
