Senior Deep Learning Software Engineer, Inference
Design and optimize GPU-accelerated deep learning inference software for large-scale language and generative AI models. Develop and enhance high-performance frameworks like vLLM and SGLang, implementing cutting-edge algorithms and performance improvements across NVIDIA's GPU architectures. Collaborate with open-source communities and internal teams to optimize model serving pipelines using tools like CUDA, Triton, and CUTLASS.