National AI Awards 2025Discover AI's trailblazers! Join us to celebrate innovation and nominate industry leaders.

Nominate & Attend

Platform engineer, MLOps (UK)

writer.com
London
1 month ago
Applications closed

Related Jobs

View all jobs

MLOps Data Platform Engineer

Senior Software Engineer - MLOps (Basé à London)

Senior Software Engineer - MLOps (Basé à London)

MLOps Field Engineer

Machine Learning Ops Engineer - AI

Automation Engineer

About this role

As a Platform engineer, MLOps, you will be critical to deploying and managing cutting-edge infrastructure crucial for AI/ML operations, and you will collaborate with AI/ML engineers and researchers to develop a robust CI/CD pipeline that supports safe and reproducible experiments. Your expertise will also extend to setting up and maintaining monitoring, logging, and alerting systems to oversee extensive training runs and client-facing APIs. You will ensure that training environments are optimally available and efficiently managed across multiple clusters, enhancing our containerization and orchestration systems with advanced tools like Docker and Kubernetes.

This role demands a proactive approach to maintaining large Kubernetes clusters, optimizing system performance, and providing operational support for our suite of software solutions. If you are driven by challenges and motivated by the continuous pursuit of innovation, this role offers the opportunity to make a significant impact in a dynamic, fast-paced environment.

????️ Your responsibilities:

  • Work closely with AI/ML engineers and researchers to design and deploy a CI/CD pipeline that ensures safe and reproducible experiments.

  • Set up and manage monitoring, logging, and alerting systems for extensive training runs and client-facing APIs.

  • Ensure training environments are consistently available and prepared across multiple clusters.

  • Develop and manage containerization and orchestration systems utilizing tools such as Docker and Kubernetes.

  • Operate and oversee large Kubernetes clusters with GPU workloads.

  • Improve reliability, quality, and time-to-market of our suite of software solutions

  • Measure and optimize system performance, with an eye toward pushing our capabilities forward, getting ahead of customer needs, and innovating for continual improvement

  • Provide primary operational support and engineering for multiple large-scale distributed software applications

️ Is this you?

  • You have professional experience with:

    • Model training

    • Huggingface Transformers

    • Pytorch

    • vLLM

    • TensorRT

    • Infrastructure as code tools like Terraform

    • Scripting languages such as Python or Bash

    • Cloud platforms such as Google Cloud, AWS or Azure

    • Git and GitHub workflows

    • Tracing and Monitoring

  • Familiar with high-performance, large-scale ML systems

  • You have a knack for troubleshooting complex systems and enjoy solving challenging problems

  • Proactive in identifying problems, performance bottlenecks, and areas for improvement

  • Take pride in building and operating scalable, reliable, secure systems

  • Are comfortable with ambiguity and rapid change

Preferred skills and experience:

  • Familiar with monitoring tools such as Prometheus, Grafana, or similar

  • 5+ years building core infrastructure

  • Experience running inference clusters at scale

  • Experience operating orchestration systems such as Kubernetes at scale

    Benefits & perks (UK full-time employees):

    • Generous PTO, plus company holidays

    • Comprehensive medical and dental insurance

    • Paid parental leave for all parents (12 weeks)

    • Fertility and family planning support

    • Early-detection cancer testingthrough Galleri

    • Competitive pension scheme and company contribution

    • Annual work-life stipends for:

      • Home office setup, cell phone, internet

      • Wellness stipend for gym, massage/chiropractor, personal training, etc.

      • Learning and development stipend

    • Company-wide off-sites and team off-sites

    • Competitive compensation and company stock options

    #LI-Remote


#J-18808-Ljbffr

National AI Awards 2025

Subscribe to Future Tech Insights for the latest jobs & insights, direct to your inbox.

By subscribing, you agree to our privacy policy and terms of service.

Industry Insights

Discover insightful articles, industry insights, expert tips, and curated resources.

How to Find Hidden Machine Learning Jobs in the UK Using Professional Bodies like BCS, Turing Society & More

Machine learning (ML) continues to transform sectors across the UK—from fintech and retail to healthtech and autonomous systems. But while the demand for ML engineers, researchers, and applied scientists is growing, many of the best opportunities are never posted on traditional job boards. So, where do you find them? The answer lies in professional bodies, academic-industry networks, and tight-knit ML communities. In this guide, we’ll show you how to uncover hidden machine learning jobs in the UK by engaging with groups like the BCS (The Chartered Institute for IT), Turing Society, Alan Turing Institute, and others. We’ll explore how to use member directories, CPD events, SIGs (Special Interest Groups), and community projects to build connections, gain early access to job leads, and raise your professional profile in the ML ecosystem.

How to Get a Better Machine Learning Job After a Lay-Off or Redundancy

Redundancy in machine learning can feel especially frustrating when your role was technically advanced, strategically important, or AI-facing. But the UK still has strong demand for machine learning professionals across fintech, healthtech, retail, cybersecurity, autonomous systems, and generative AI. Whether you're a research-oriented ML engineer, production-focused MLOps developer, or applied scientist, this guide is designed to help you bounce back from redundancy and find a better opportunity that suits your goals.

Machine Learning Jobs Salary Calculator 2025: Figure Out Your True Worth in Seconds

Why last year’s pay survey is useless for UK ML professionals today Ask a Machine Learning Engineer wrangling transformer checkpoints, an MLOps Lead firefighting drift alarms, or a Research Scientist training diffusion models at 3 a.m.: “Am I earning what I deserve?” The honest answer changes monthly. A single OpenAI model drop doubles GPU demand, healthcare regulators release fresh explainability guidance, & a fintech unicorn pays six figures for vector‑search expertise. Each shock nudges salary bands. Any PDF salary guide printed in 2024 now looks like an outdated Jupyter notebook—missing the gen‑AI tsunami, the surge in edge inference, & the UK’s new Responsible‑AI framework. To give ML professionals an accurate benchmark, MachineLearningJobs.co.uk distilled a transparent, three‑factor formula that estimates a realistic 2025 salary in under a minute. Feed in your discipline, UK region, & seniority; you’ll receive a defensible figure—no stale averages, no guesswork. This article unpacks the formula, highlights the forces driving ML pay skyward, & offers five practical moves to boost your value inside the next ninety days.