Staff Software Engineer, AI Reliability Engineering
This role focuses on ensuring the reliability and resilience of large language model serving systems across the full stack, from API layers to accelerators. You'll design service level objectives, build observability systems, lead incident response, and collaborate across teams to strengthen critical infrastructure. The position demands a systems thinker with strong distributed systems experience and a proactive approach to improving robustness in high-scale AI environments.