Senior Software Engineer, AI Inference Systems
Design and optimize high-performance AI inference systems for large-scale models, focusing on GPU kernel development, compiler optimization, and distributed inference frameworks. Work on vLLM, speculative decoding, and MLPerf benchmarking while collaborating across compiler, scheduling, and performance teams. Contribute to open-source and publish research to advance ML systems.