Senior Software Engineer, AI Inference Systems
This role involves building and optimizing high-performance AI inference systems for large-scale models on NVIDIA GPUs. You'll contribute to open-source frameworks like vLLM, optimize GPU kernels and compilers, design benchmarking methodologies, and work on scheduling for multi-node, multi-cloud deployments. The position blends deep systems engineering, performance optimization, and research to advance the state of accelerated AI computing.