Senior Deep Learning Software Engineer, Inference
This role involves designing, optimizing, and maintaining high-performance deep learning inference software for large-scale AI models, particularly focusing on GPU-accelerated frameworks like vLLM and SGLang. The engineer will implement cutting-edge algorithms, improve model serving efficiency across NVIDIA's GPU architectures, and contribute to open-source inference libraries. Work includes performance tuning, cross-architecture scaling, and close collaboration with deep learning and framework teams.