Research Engineer - Inference
The role involves deploying and optimizing cutting-edge AI models for real-time inference, focusing on performance, reliability, and scalability. You'll work across the stack to improve latency, throughput, and cost efficiency, building high-performance serving systems and tooling that bridge research and production. The position emphasizes autonomous problem-solving and deep optimization in GPU and ML infrastructure environments.