Senior Solutions Architect – Large Scale AI Inference
This role involves guiding AI-native customers in deploying and optimizing large-scale inference workloads on multi-node GPU clusters, with a focus on architecting efficient pipelines for dense and sparse Mixture-of-Experts models. The Senior Solutions Architect will tackle challenges in inference efficiency, including quantization, speculative decoding, and KV cache management, while collaborating with NVIDIA product teams and engaging the EMEA developer community through technical leadership and knowledge sharing.