Senior Software Engineer, AI Inference Systems
This role involves building and optimizing AI inference systems for large-scale models on NVIDIA GPUs. The engineer will work on high-performance inference stacks, GPU kernel optimization, compiler infrastructure, and distributed scheduling, while contributing to benchmarks like MLPerf. Collaboration with research teams and integration of cutting-edge ML systems concepts into production software are key aspects of the role.