Senior HPC AI Cluster Engineer
This role involves designing, implementing, and maintaining large-scale HPC and AI clusters, with a focus on GPU-accelerated computing and deep learning platforms. The engineer will automate infrastructure deployment, manage orchestration tools like Slurm and Kubernetes, and support R&D through proof-of-concept projects. Work includes full-stack troubleshooting, performance tuning, and developing self-service solutions for scientific and technical users.