JobsAI Infrastructure Engineer
Job description
The AI Infrastructure Engineer role focuses on optimizing LLM inference on Intel's next-generation GPU architectures. The engineer will work across the entire inference stack, profiling performance bottlenecks and developing custom GPU kernels. This position requires collaboration with hardware teams and contributions to open-source projects. The ideal candidate will have a passion for AI infrastructure and performance optimization.
Requirements
- Bachelor's Degree in Computer Science, Software Engineering, Artificial Intelligence/Machine Learning, or related field and 4+ years experience, Masters Degree and 3+ years, OR PhD.
- 3+ years of relevant software engineering experience in GPU computing, AI systems, or high-performance computing (HPC).
- Proficiency in modern C++ and Python.
Responsibilities
- Own the end-to-end optimization pipeline for running state-of-the-art LLMs on Intel GPUs.
- Profile, diagnose, and resolve cross-stack performance bottlenecks.
- Design, write, and optimize custom high-performance kernels for critical attention mechanisms, MoE, quantization, and operator fusions.
- Upstream architectural improvements and hardware backends directly into open-source repositories like vLLM, SGLang, and PyTorch.
- Apply roofline analysis and systematic profiling to decompose bottlenecks and shape future GPU roadmaps.
Benefits
- Intel offers a comprehensive benefits package including competitive pay, stock programs, healthcare coverage, retirement plans, paid time off, parental leave, and programs supporting employee wellbeing and professional development.
Is this posting expired or inaccurate?
