ML Infrastructure Engineer Jobs

Engineers who design, build, and maintain the systems that power machine learning workflows. A critical role in scaling and optimising ML operations.

Open roles
10
Salary range
£119k – £457k
Hiring companies
5

ML Infrastructure Engineers are the backbone of any machine learning organisation. They design, build, and maintain the infrastructure that supports the entire ML lifecycle, from data ingestion and preprocessing to model training, deployment, and monitoring. These engineers work closely with data scientists and ML researchers to ensure that the infrastructure is robust, scalable, and efficient. They are often found in tech scaleups, research-heavy startups, and the larger consultancies, where the demand for reliable and performant ML systems is high.

What the role does

Inside the role of an ML Infrastructure Engineer

A typical week is split between designing and implementing new infrastructure components, troubleshooting existing systems, and collaborating with cross-functional teams.

  1. 01
    Design and implement scalable data pipelines.
  2. 02
    Optimise and maintain ML model training environments.
  3. 03
    Collaborate with data scientists to understand infrastructure needs.
  4. 04
    Troubleshoot and resolve performance issues in production systems.
  5. 05
    Document and standardise infrastructure best practices.
  6. 06
    Stay updated with the latest ML infrastructure trends and tools.
Salary on the board

£119k – £457k

Based on advertised midpoints across the 4 priced listings posted in the last 12 months. Base salary only.

Salary visibility
29% of listings advertise a salary.
Skills & tools

What hiring managers ask for

% of 6 listings posted in the last 12 months that mention each skill, extracted from job descriptions.

GCP
83%
AWS
83%
CI/CD
83%
Python
67%
Azure
67%
Terraform
67%
Kubernetes
67%
Networking
33%
Linux
33%
PyTorch
17%
Distributed Training
17%
Docker
17%
Career ladder

From Junior to Principal

A typical UK progression for ml infrastructure engineers. Years are guidance — strong people move faster, and many senior folks sidestep into research, product or management.

  1. Level 1

    Junior ML Infrastructure Engineer

    0–2 yrs

    Assist in the design and implementation of basic ML infrastructure components, with a focus on learning and gaining hands-on experience.

  2. Level 2

    ML Infrastructure Engineer

    2–5 yrs

    Own the design and implementation of key infrastructure components, ensuring they are scalable and efficient.

  3. Level 3

    Senior ML Infrastructure Engineer

    5–8 yrs

    Lead the development of complex infrastructure solutions, mentor junior engineers, and drive infrastructure innovation.

  4. Level 4

    Principal ML Infrastructure Engineer

    8+ yrs

    Strategise and oversee the entire ML infrastructure architecture, influence company-wide technical decisions, and lead major initiatives.

Pathway

How to become a ML Infrastructure Engineer

There's no single route, but most people follow some version of these steps.

  1. 1

    Learn the Basics

    Start by gaining a solid understanding of data engineering, cloud services, and ML workflows. Hands-on projects and courses can help build a strong foundation.

  2. 2

    Gain Practical Experience

    Work on real-world projects, either through internships or entry-level roles. Focus on building and maintaining data pipelines and ML environments.

  3. 3

    Specialise in ML Infrastructure

    Deepen your expertise in specific areas such as distributed systems, containerisation, and orchestration tools. Contribute to more complex infrastructure designs.

  4. 4

    Lead Projects and Teams

    Take on leadership roles, managing infrastructure projects and mentoring junior engineers. Drive innovation and best practices within your team.

  5. 5

    Influence Company Strategy

    At the senior level, you will influence the overall ML infrastructure strategy, working closely with leadership to shape the technical direction of the organisation.

  6. 6

    Industry Thought Leadership

    Becoming a principal engineer involves not only leading within your organisation but also contributing to the broader ML community through speaking engagements and publications.

Live jobs

10 live roles

See all 10 roles

Staff Infrastructure Engineer, Cluster Infrastructure

This role involves leading the technical strategy for agent-driven automation in cluster lifecycle management, ensuring secure, scalable, and fault-tolerant compute infrastructure across cloud and on-prem environments. You'll collaborate with research, product, and security teams to shape long-term infrastructure direction, with a focus on high-bandwidth interconnectivity and operational excellence. The position emphasizes mentorship, cross-team alignment, and driving innovation in large-scale cluster provisioning and management.

Anthropic London, United Kingdom £325,000 – £485,000 pa
On-site Permanent
PhysicsX logo

Senior Machine Learning Infrastructure Engineer, Research

This role involves designing and operating distributed training infrastructure for large physics models using NVIDIA DGX B200 systems, with a focus on optimizing training pipelines, data I/O performance, and model serving. The engineer will work closely with research scientists and ML engineers to enable efficient, scalable AI training and deployment in high-performance computing environments, while also building observability and reproducibility into the research workflow.

PhysicsX United Kingdom
PhysicsX logo

Principal Machine Learning Infrastructure Engineer

This role focuses on designing and operating scalable machine learning infrastructure for training and serving large physics-based models. You'll work closely with research scientists and ML engineers to optimize distributed training pipelines, improve data I/O performance, and build reliable model serving systems. The position emphasizes systems-level problem-solving, infrastructure automation, and enabling fast, reproducible experimentation on high-performance GPU clusters.

PhysicsX London, United Kingdom
OpenAI logo

Software Engineer, ChatGPT Infrastructure

Design, build, and operate reliable, scalable systems supporting AI products like ChatGPT and the OpenAI API. Focus on performance optimization, system resilience, automation, and developer experience across distributed infrastructure. Collaborate with research, product, and engineering teams to improve reliability, scalability, and observability at global scale.

OpenAI London, United Kingdom
On-site Permanent
OpenAI logo

Software Engineer, Cloud Infrastructure

Design and build scalable, secure cloud infrastructure platforms that support OpenAI's products, including ChatGPT and the API. Work with Kubernetes, networking, and cloud abstractions while ensuring system reliability and contributing to a safety-first culture. Participate in on-call rotations and help scale systems to meet growing demand.

OpenAI London, United Kingdom
Permanent

Staff Software Engineer, Infrastructure (Distributed Systems)

Lead and architect large-scale distributed systems that support AI model training, serving, and security. Work closely with research and product teams to build reliable, scalable infrastructure on Kubernetes and cloud platforms. Drive technical strategy, mentor engineers, and improve operational resilience across high-impact systems.

Anthropic London, United Kingdom £325,000 – £390,000 pa

Forward Deployed Engineer, Infrastructure Specialist (Europe/Middle East)

This role involves leading end-to-end deployments of Cohere’s North AI platform in private cloud and on-premises environments for enterprise clients. You’ll collaborate closely with client engineering and IT teams to design secure, compliant deployment strategies, troubleshoot issues, and ensure seamless integration with existing workflows in sectors like finance and healthcare. The role emphasizes hands-on technical execution, Kubernetes expertise, and adapting rapidly to evolving customer requirements in high-stakes environments.

Cohere United Kingdom
Remote Permanent

Research Infrasructure Engineer

Research Infrastructure EngineerLocation: Hampshire, South Coast / Hybrid Working AvailableThe OpportunityWe're supporting a leading research-led organisation in the search for a Research Infrastructure Engineer to join a specialist team delivering advanced computing platforms that support research, data science and AI...

Tate London City Southampton, United Kingdom £41,000 – £49,000 pa
Top hirers

Companies hiring ml infrastructure engineers

See all companies →
Hiring locations

Where this role is hiring

The locations with the most live listings for this role today.

FAQs

Common questions

  • Essential skills include strong programming abilities, knowledge of cloud platforms, experience with data pipelines, and a deep understanding of machine learning workflows.

  • Gain experience in data engineering and cloud services, and start working on ML-related projects. Online courses and certifications can also help bridge the gap.

  • Senior engineers lead the design and implementation of complex infrastructure solutions, mentor junior team members, and drive innovation in infrastructure practices.

  • The typical progression is from Junior to Senior, then to Principal Engineer. Each step involves increasing responsibility and influence over the organisation's technical direction.

  • Salaries vary based on experience and location. For specific salary ranges, please refer to the salary section on this page.

Hiring ml infrastructure engineers?

Post your role in 90 seconds and reach the specialist audience that already reads this page.