ML Infrastructure Engineer Jobs

Engineers who design, build, and maintain the systems that power machine learning workflows. A critical role in scaling and optimising ML operations.

Open roles
8
Salary range
£43k – £93k
Hiring companies
6

ML Infrastructure Engineers are the backbone of any machine learning organisation. They design, build, and maintain the infrastructure that supports the entire ML lifecycle, from data ingestion and preprocessing to model training, deployment, and monitoring. These engineers work closely with data scientists and ML researchers to ensure that the infrastructure is robust, scalable, and efficient. They are often found in tech scaleups, research-heavy startups, and the larger consultancies, where the demand for reliable and performant ML systems is high.

What the role does

Inside the role of an ML Infrastructure Engineer

A typical week is split between designing and implementing new infrastructure components, troubleshooting existing systems, and collaborating with cross-functional teams.

  1. 01
    Design and implement scalable data pipelines.
  2. 02
    Optimise and maintain ML model training environments.
  3. 03
    Collaborate with data scientists to understand infrastructure needs.
  4. 04
    Troubleshoot and resolve performance issues in production systems.
  5. 05
    Document and standardise infrastructure best practices.
  6. 06
    Stay updated with the latest ML infrastructure trends and tools.
Salary on the board

£43k – £93k

Based on advertised midpoints across the 3 priced listings posted in the last 12 months. Base salary only.

Salary visibility
23% of listings advertise a salary.
Skills & tools

What hiring managers ask for

% of 7 listings posted in the last 12 months that mention each skill, extracted from job descriptions.

Python
71%
GCP
71%
AWS
71%
CI/CD
71%
Kubernetes
71%
Azure
57%
Terraform
57%
Linux
43%
Networking
29%
Go
29%
Rust
29%
Distributed Systems
29%
Career ladder

From Junior to Principal

A typical UK progression for ml infrastructure engineers. Years are guidance — strong people move faster, and many senior folks sidestep into research, product or management.

  1. Level 1

    Junior ML Infrastructure Engineer

    0–2 yrs

    Assist in the design and implementation of basic ML infrastructure components, with a focus on learning and gaining hands-on experience.

  2. Level 2

    ML Infrastructure Engineer

    2–5 yrs

    Own the design and implementation of key infrastructure components, ensuring they are scalable and efficient.

  3. Level 3

    Senior ML Infrastructure Engineer

    5–8 yrs

    Lead the development of complex infrastructure solutions, mentor junior engineers, and drive infrastructure innovation.

  4. Level 4

    Principal ML Infrastructure Engineer

    8+ yrs

    Strategise and oversee the entire ML infrastructure architecture, influence company-wide technical decisions, and lead major initiatives.

Pathway

How to become a ML Infrastructure Engineer

There's no single route, but most people follow some version of these steps.

  1. 1

    Learn the Basics

    Start by gaining a solid understanding of data engineering, cloud services, and ML workflows. Hands-on projects and courses can help build a strong foundation.

  2. 2

    Gain Practical Experience

    Work on real-world projects, either through internships or entry-level roles. Focus on building and maintaining data pipelines and ML environments.

  3. 3

    Specialise in ML Infrastructure

    Deepen your expertise in specific areas such as distributed systems, containerisation, and orchestration tools. Contribute to more complex infrastructure designs.

  4. 4

    Lead Projects and Teams

    Take on leadership roles, managing infrastructure projects and mentoring junior engineers. Drive innovation and best practices within your team.

  5. 5

    Influence Company Strategy

    At the senior level, you will influence the overall ML infrastructure strategy, working closely with leadership to shape the technical direction of the organisation.

  6. 6

    Industry Thought Leadership

    Becoming a principal engineer involves not only leading within your organisation but also contributing to the broader ML community through speaking engagements and publications.

Live jobs

8 live roles

Synthesia logo

Infrastructure Engineer

This role involves maintaining and scaling Kubernetes clusters, managing AWS and GCP cloud environments, and improving CI/CD systems. You will also focus on observability, FinOps practices, and collaborating with product engineers to deploy and monitor production services.

Synthesia London, United Kingdom
Remote Permanent
PhysicsX logo

Senior Infrastructure Engineer, Research

This role involves designing and operating distributed training infrastructure for large physics models using NVIDIA DGX B200 systems, with a focus on optimizing training pipelines, data I/O performance, and model serving. The engineer will work closely with research scientists and ML engineers to enable efficient, scalable AI training and deployment in high-performance computing environments, while also building observability and reproducibility into the research workflow.

PhysicsX United Kingdom

Staff+ Infrastructure Engineer, Cluster Infrastructure

This role involves leading the technical strategy for agent-driven automation in cluster lifecycle management, ensuring secure, scalable, and fault-tolerant compute infrastructure across cloud and on-prem environments. You'll collaborate with research, product, and security teams to shape long-term infrastructure direction, with a focus on high-bandwidth interconnectivity and operational excellence. The position emphasizes mentorship, cross-team alignment, and driving innovation in large-scale cluster provisioning and management.

Anthropic London, United Kingdom £325,000 – £485,000 pa
On-site Permanent
OpenAI logo

Software Engineer, GPU Infrastructure- ChatGPT Engineering

Design and operate software systems managing large-scale GPU clusters that power ChatGPT inference workloads. Build automation, observability, and intelligent tooling to improve fleet reliability, efficiency, and scalability. Collaborate with infrastructure, research, and product teams to optimize compute utilization and drive operational excellence across distributed systems.

OpenAI London, United Kingdom
Hybrid Permanent
OpenAI logo

Software Engineer, ChatGPT Infrastructure

Design, build, and operate reliable, scalable systems supporting AI products like ChatGPT and the OpenAI API. Focus on performance optimization, system resilience, automation, and developer experience across distributed infrastructure. Collaborate with research, product, and engineering teams to improve reliability, scalability, and observability at global scale.

OpenAI London, United Kingdom
On-site Permanent

Staff+ Software Engineer, Safeguards Infrastructure

Build foundational systems for AI safety, oversight, and intervention mechanisms, focusing on detecting unwanted model behaviors and preventing disallowed use. Develop infrastructure for data management, metrics, evaluations, and review tooling while ensuring high operational reliability and scalability. Work across the stack to create multi-layered, real-time safety defenses with minimal human intervention.

Anthropic London, United Kingdom £255,000 – £325,000 pa
Hybrid Permanent

Research Infrasructure Engineer

Research Infrastructure EngineerLocation: Hampshire, South Coast / Hybrid Working AvailableThe OpportunityWe're supporting a leading research-led organisation in the search for a Research Infrastructure Engineer to join a specialist team delivering advanced computing platforms that support research, data science and AI...

Tate London City Southampton, United Kingdom £41,000 – £49,000 pa

Senior Research Infrasructure Engineer

Senior Research Infrastructure EngineerLocation: Hampshire, South Coast / Hybrid Working AvailableThe OpportunityWe're supporting a leading research-led organisation in the search for a Senior Research Infrastructure Engineer to play a key role in the development and delivery of advanced research computing...

Tate London City Southampton, United Kingdom £51,000 – £63,000 pa
Hiring locations

Where this role is hiring

The locations with the most live listings for this role today.

FAQs

Common questions

  • Essential skills include strong programming abilities, knowledge of cloud platforms, experience with data pipelines, and a deep understanding of machine learning workflows.

  • Gain experience in data engineering and cloud services, and start working on ML-related projects. Online courses and certifications can also help bridge the gap.

  • Senior engineers lead the design and implementation of complex infrastructure solutions, mentor junior team members, and drive innovation in infrastructure practices.

  • The typical progression is from Junior to Senior, then to Principal Engineer. Each step involves increasing responsibility and influence over the organisation's technical direction.

  • Salaries vary based on experience and location. For specific salary ranges, please refer to the salary section on this page.

Hiring ml infrastructure engineers?

Post your role in 90 seconds and reach the specialist audience that already reads this page.