Senior Solutions Architect – Large Scale AI Inference
This role involves guiding AI customers in deploying and optimizing large-scale inference workloads on multi-node GPU clusters, with a focus on advanced techniques like MoE serving, speculative decoding, and memory optimization. The architect will design efficient inference pipelines for dense and sparse models across thousands of GPUs and collaborate with NVIDIA’s product teams to enhance customer success. The position also includes leading technical engagement through workshops and reference architectures across EMEA.