Staff SRE, AI Infrastructure
This role involves building and scaling the reliability foundations of a large-scale AI and GPU compute platform. You will define SRE frameworks, automation, and operational standards for model development and training infrastructure, ensuring high availability and performance. The position bridges AI research, cloud infrastructure, and production operations, with a focus on observability, incident response, and resilient system design.