Senior Deep Learning Software Engineer, Inference
Design and optimize high-performance deep learning inference software for NVIDIA's GPU-accelerated platforms, focusing on frameworks like vLLM and SGLang. Work on performance tuning of LLMs and generative AI models across datacenter and edge GPUs. Collaborate with open-source communities and internal teams to enhance model serving efficiency and scalability.