Senior Machine Learning Infrastructure Engineer, Research
This role involves designing and operating distributed training infrastructure for large physics models using NVIDIA DGX B200 systems, with a focus on optimizing training pipelines, data I/O performance, and model serving. The engineer will work closely with research scientists and ML engineers to enable efficient, scalable AI training and deployment in high-performance computing environments, while also building observability and reproducibility into the research workflow.