Staff ML Performance Engineer (Inference Optimisation)
Optimise machine learning inference performance for edge and GPU accelerators, focusing on efficient execution of large transformer models in low-power, cost-constrained environments. Work across the full stack—from model graphs and compilers to kernels and embedded deployment—to deliver measurable improvements in latency, memory, and power. Collaborate with model developers and contribute to tooling and technical roadmaps for production-grade autonomous driving systems.