Senior Solutions Architect – Large Scale Neural Networks Inference
Lead technical strategy for large-scale AI inference deployments across EMEA, working closely with frontier AI labs and enterprises. Architect and optimize high-performance inference pipelines using NVIDIA's stack, including TensorRT-LLm, vLLM, and Dynamo, while translating real-world deployment challenges into product improvements. Drive end-to-end engagements from proof of concept to production at scale, focusing on efficiency, latency, and GPU utilization.