Principal Machine Learning Infrastructure Engineer
This role focuses on designing and operating scalable machine learning infrastructure for training and serving large physics-based models. You'll work closely with research scientists and ML engineers to optimize distributed training pipelines, improve data I/O performance, and build reliable model serving systems. The position emphasizes systems-level problem-solving, infrastructure automation, and enabling fast, reproducible experimentation on high-performance GPU clusters.