Latest Cloud Infrastructure Jobs

NVIDIA logo

Senior Cloud Infrastructure and DevOps Solutions Architect

This role involves owning full-solution validation, minimizing time to production, and ensuring Day 2 stability for large-scale GPU clusters. You will work closely with customers, partners, and cross-functional teams to architect and implement Kubernetes-based platforms, automate processes, and provide consultative guidance across the entire stack.

NVIDIA

Staff Software Engineer, Inference

This role involves building and maintaining large-scale, compute-agnostic inference systems that power Claude for millions of users and support cutting-edge AI research. The engineer will work across the full stack, optimizing intelligent request routing, fleet orchestration, and performance across diverse AI accelerators in multi-cloud environments. Projects include autoscaling infrastructure, integrating new hardware, and developing production pipelines for model deployment.

Anthropic London, United Kingdom £325,000 – £390,000 pa
Hybrid Permanent
Wayve logo

Platform Engineer, AI Enablement

This role involves building and operating secure, scalable infrastructure to enable safe and cost-effective access to AI models and agentic workflows across the company. You'll design governed model-access layers, implement observability and safety controls, and create reusable platform primitives. The position requires close collaboration with security, IT, and engineering teams to support AI adoption while ensuring compliance, reliability, and operational excellence.

Wayve London, United Kingdom

Engineering - Internal AI Transformation

As an Internal AI Engineer at ElevenLabs, you will design and implement AI-driven workflows to automate manual processes across various teams, including GTM, Operations, and Finance. You will work closely with business owners to identify pain points, build solutions, and ensure the reliability and safety of AI systems.

ElevenLabs United Kingdom
Remote Permanent
PhysicsX logo

Senior Machine Learning Infrastructure Engineer, Research

This role involves designing and operating distributed training infrastructure for large physics models using NVIDIA DGX B200 systems, with a focus on optimizing training pipelines, data I/O performance, and model serving. The engineer will work closely with research scientists and ML engineers to enable efficient, scalable AI training and deployment in high-performance computing environments, while also building observability and reproducibility into the research workflow.

PhysicsX United Kingdom
Wayve logo

Software Engineer, AI Libraries

Design and build scalable Python libraries and tools that support machine learning workflows for autonomous driving systems. Focus on creating reliable, modular abstractions for data loading, distributed training, and model evaluation used by ML engineers and researchers. Work closely with ML teams to improve the performance and usability of training infrastructure across large GPU clusters and cloud environments.

Wayve London, United Kingdom

Senior Software Engineer

This role involves designing and building scalable backend services, developing low-latency and high-throughput systems, and solving complex scalability and concurrency problems. You'll work with Python, Java, and cloud-native infrastructure, contributing to system architecture and technical direction in a quantitative sports betting environment.

Method Resourcing London, United Kingdom £120,000 pa

Software Engineer, Python

Develop and optimize AI-driven video processing systems using Python and related technologies, focusing on scalable media delivery and low-latency streaming. Work across cloud-native infrastructure and immersive content platforms, integrating AI and video processing pipelines. Collaborate on real-time and batch processing solutions with exposure to GPU-accelerated frameworks and containerized environments.

Enterprise Recruitment Sheldon Square, London, United Kingdom £60,000 – £90,000 pa
Faculty AI logo

Machine Learning Engineer

Develop and deploy production-grade machine learning systems for high-impact clients, particularly in the defence sector. Design scalable ML infrastructure, lead technical architecture decisions, and translate complex AI concepts for stakeholders. Work across the full ML lifecycle using cloud platforms and containerisation technologies.

Faculty AI London, United Kingdom
Hybrid Permanent Flexible Clearance Required

Machine Learning Engineer (Safety)

As a Machine Learning Engineer, you will build and deploy production-grade ML systems, collaborate with cross-functional teams to solve critical client challenges, and lead technical scoping and architectural decisions. You will work on high-stakes, high-impact missions in national security and AI safety, ensuring AI is secure, trustworthy, and safe for all.

Faculty London, United Kingdom
Hybrid Permanent

Forward Deployed Engineer

This role involves building and deploying production-grade machine learning systems for high-impact clients in national security and AI safety. You'll work across the full ML lifecycle, design scalable software architecture, and act as a technical advisor to translate complex concepts for stakeholders. The role emphasizes operationalising models, defining deployment standards, and delivering bespoke AI solutions in collaboration with cross-functional teams.

Faculty London, United Kingdom
Permanent

Senior Machine Learning Research Engineer

This role involves building and scaling generative and predictive machine learning models for cellular behaviour using multi-omic and interventional biological data. The engineer will implement large neural networks, optimise distributed training and inference pipelines, and collaborate closely with ML scientists and biologists to translate research into production-grade systems. Work focuses on robust infrastructure development, performance optimisation, and integrating cutting-edge ML tooling within an interdisciplinary team.

Relation Therapeutics London, United Kingdom
Permanent

Staff Infrastructure Engineer, Cluster Infrastructure

This role involves leading the technical strategy for agent-driven automation in cluster lifecycle management, ensuring secure, scalable, and fault-tolerant compute infrastructure across cloud and on-prem environments. You'll collaborate with research, product, and security teams to shape long-term infrastructure direction, with a focus on high-bandwidth interconnectivity and operational excellence. The position emphasizes mentorship, cross-team alignment, and driving innovation in large-scale cluster provisioning and management.

Anthropic London, United Kingdom £325,000 – £485,000 pa
On-site Permanent

Public Cloud - Forward Deployed Engineer VP - Citi

As a Forward Deployed Engineer VP, you will contribute to Citi's public cloud strategy by managing container fleet services, automating processes, and collaborating with cross-functional teams. You will work on innovative projects like HPC platforms and GenAI, using technologies such as AWS, GCP, Python, and Kubernetes.

eFinancialCareers Belfast, United Kingdom

AWS Cloud DevOps Engineer

This role involves architecting, deploying, and supporting AI and machine learning solutions on AWS. You will design ML pipelines, develop scalable tools, and facilitate the deployment of proof-of-concept systems, working closely with data scientists and clients to ensure auditability and security.

eFinancialCareers London, United Kingdom