Latest infrastructure engineer Jobs

NVIDIA logo

Senior HPC AI Cluster Engineer

Design and maintain large-scale HPC/AI clusters with a focus on automation, monitoring, and performance optimization across bare metal, OS, and application layers. Develop CI/CD pipelines and self-service tooling for infrastructure management, supporting cutting-edge research and development in AI and accelerated computing. Collaborate with specialists in GPU compute, networking, and storage to deploy and troubleshoot high-performance systems at scale.

NVIDIA
Remote Permanent
NVIDIA logo

Senior HPC AI Cluster Engineer

This role involves designing, implementing, and maintaining large-scale HPC/AI clusters, managing job/workload schedules, and developing CI/CD pipelines. You will work closely with HPC, OS, GPU compute, and systems specialists to architect and optimize performance platforms, and support R&D activities.

NVIDIA
Wayve logo

Staff ML Engineer Gaia

As a Staff ML Engineer on Gaia, you will lead and execute large-scale training runs for video foundation models, contribute to model architecture and training strategies, and improve world-model capabilities for synthetic scenario generation. You will work closely with research, applications, simulation engineering, and cloud/infrastructure teams to deliver end-to-end impact in a fast-paced, results-focused environment.

Wayve London, United Kingdom
Wayve logo

Machine Learning Engineer, Performance Tooling

Builds performance analysis tools to optimize AI models across the full stack, from compilers to hardware. Measures, predicts, and advises on latency, throughput, and compute efficiency using profiling data. Works cross-functionally with model, runtime, and hardware teams to drive data-informed performance decisions.

Wayve United Kingdom
PhysicsX logo

Senior AI / Agentic Engineer

This role involves architecting and leading the development of a production-grade agentic AI platform that integrates with engineering simulation workflows. You'll design secure, observable, and governed agent systems using durable execution, sandboxing, and evaluation frameworks, enabling both internal teams and customers to build advanced AI-driven engineering solutions. The position emphasizes robust platform design, developer experience, and real-world impact in high-fidelity simulation environments.

PhysicsX London, United Kingdom
Hybrid Permanent
Synthesia logo

Principal ML Platform Engineer

Design and build scalable, reliable systems for training, serving, and operating generative AI models in production. Develop internal tooling and agentic workflows to reduce manual effort and improve automation across research and product teams. Collaborate with researchers and engineers to enhance platform observability, debugging, and developer experience at scale.

Synthesia London, United Kingdom
Remote Permanent
PhysicsX logo

Forward Deployed Software Engineer

A Forward Deployed Software Engineer at PhysicsX works directly with customers and engineering teams to build and deploy AI-driven simulation applications that solve complex engineering problems across industries like aerospace, energy, and semiconductors. The role involves full-stack development, rapid prototyping, and on-site collaboration, with a focus on delivering production-grade tools using physics AI models. Engineers contribute to backend services, APIs, and platform features while shaping product direction based on real-world customer needs.

PhysicsX United Kingdom
Hybrid Permanent

Senior Machine Learning Research Engineer

This role involves building and scaling generative and predictive machine learning models for cellular behaviour using multi-omic and interventional biological data. The engineer will implement large neural networks, optimise distributed training and inference pipelines, and collaborate closely with ML scientists and biologists to translate research into production-grade systems. Work focuses on robust infrastructure development, performance optimisation, and integrating cutting-edge ML tooling within an interdisciplinary team.

Relation Therapeutics London, United Kingdom
Permanent

Senior Forward Deployed Engineer

Lead the design and deployment of scalable, production-grade machine learning systems for high-impact clients in national security and AI safety. Collaborate across engineering, data science, and commercial teams to operationalise AI solutions, while mentoring junior engineers and shaping technical standards. Work on-site with government clients and contribute to ethical, secure AI practices in a cross-functional, client-facing role.

Faculty London, United Kingdom
Permanent

Full-Stack Engineer - Creative Studio

Develop and maintain full-stack components for ElevenLabs' Creative Studio products, including image/video, music, and generative AI features. Work across API and UI layers to build new proof-of-concept products and scale them, while collaborating closely with engineering, growth, and sales teams. Focus on high-ownership projects involving cutting-edge voice models and large-scale data systems.

ElevenLabs United Kingdom
Remote Permanent

Senior Machine Learning Engineer

Senior Machine Learning EngineerSalary: £80,000 - £90,000Location: London|HybridData Idols are working with an innovative technology business that is continuing to invest in its Machine Learning and AI capabilities. They are looking for a Senior Machine Learning Engineer to help build,...

Data Idols London, United Kingdom £80,000 – £90,000 pa

Senior ML Systems Engineer, Frameworks & Tooling

Design and maintain core components of a large-scale LLM training framework, focusing on distributed systems, performance optimization, and developer tooling. Work across the ML stack to improve throughput, stability, and reproducibility on multi-node GPU clusters. Build monitoring, debugging, and automation tools to support fast-moving research and production workflows.

Cohere London, United Kingdom
Remote Permanent

Engineering - Internal AI Transformation

About ElevenLabsElevenLabs is an AI research and product company transforming how we interact with technology.We launched in January 2023 with the first human-like AI voice model. Today, we serve millions of users and thousands of businesses - from fast-growing startups...

ElevenLabs United Kingdom
Remote Permanent
Ocado logo

Senior Machine Learning Engineer (E3)

This role involves owning the full machine learning lifecycle for production systems that enhance customer experience in online grocery, including personalisation, recommendations, and search ranking. The engineer will work within a cross-functional data science team to build scalable, reusable ML components while championing data quality, MLOps best practices, and platform thinking. A strong focus is placed on system reliability, technical leadership, and collaboration across disciplines.

Ocado United Kingdom
Hybrid Permanent

Member of Technical Staff, Training Performance Engineer

This role involves optimizing the performance of large language models during training, focusing on improving throughput and accelerator utilization. You'll design high-performance software, write low-level CUDA and Triton kernels, and develop profiling tools to eliminate bottlenecks. The position sits within a research-driven team working at the intersection of ML and systems engineering, leveraging large-scale infrastructure to advance NLP capabilities.

Cohere London, United Kingdom
Hybrid Permanent