Senior Software Engineer, AI Inference Systems
Design and optimize high-performance AI inference systems for large-scale models on NVIDIA GPUs, focusing on frameworks like vLLM, kernel optimization, compiler infrastructure, and distributed execution. Develop benchmarking methodologies, contribute to MLPerf, and integrate cutting-edge research into production software. Work across compiler, scheduling, and performance teams to push the limits of accelerated computing in multi-GPU and multi-cloud environments.