Senior Research Scientist, Model Evaluation
This role involves creating new evaluation benchmarks for large language models (LLMs), working with cross-functional teams to refine evaluation methods, and building scalable tools to measure model performance. The focus is on advancing the state-of-the-art in LLM evaluation and ensuring rigorous, repeatable assessments.