Senior Software Engineer, AI Inference Systems
This role involves building and optimizing AI inference systems for large-scale models, focusing on high-performance GPU stacks, kernel optimization, and multi-node deployment. The engineer will contribute to frameworks like vLLM, develop compiler infrastructure, lead benchmarking efforts including MLPerf, and integrate cutting-edge research into production software. Collaboration across compiler, scheduling, and performance teams is central to advancing NVIDIA’s accelerated computing platform.