Senior HPC AI Cluster Engineer
Design and maintain large-scale HPC/AI clusters with a focus on automation, monitoring, and performance optimization across bare metal, OS, and application layers. Develop CI/CD pipelines and self-service tooling for infrastructure management, supporting cutting-edge research and development in AI and accelerated computing. Collaborate with specialists in GPU compute, networking, and storage to deploy and troubleshoot high-performance systems at scale.