Senior HPC Performance Engineer
Conduct performance analysis and optimization of GPU communication libraries on large multi-GPU and multi-node HPC clusters. Work across the full hardware-software stack to benchmark, triage, and improve performance for deep learning and HPC applications. Develop tools to analyze performance data and collaborate with distributed teams to shape future library roadmaps.