Senior Solutions Architect – Large Scale AI Inference
This role involves guiding AI-native customers and enterprises in deploying and optimizing large-scale AI inference workloads on multi-node GPU clusters. The focus is on architecting efficient inference pipelines for dense and sparse Mixture-of-Experts (MoE) models, optimizing performance through techniques like quantization, speculative decoding, and KV cache management. The position requires deep technical collaboration with both customers and NVIDIA's product teams to advance high-performance inference solutions across EMEA.