Principal Software Reliability Engineer - Consumer Identity

Entrust

London, United Kingdom

Last month

Applications closed

Related Jobs

View all jobs

Spotlight

Senior ML Compiler Engineer

Fractile Bristol, United Kingdom

Spotlight

Machine Learning Engineer - National Security (Gloucestershire)

Mind Foundry Gloucester, Gloucestershire, United Kingdom

On-site Clearance Required

Senior Software Engineer, ML Ops

Isomorphic Labs London, United Kingdom

On-site

Senior Software Engineer Delivery Lead – Tegra System Software

NVIDIA Cambridge, United Kingdom

Senior Software Engineer Delivery Lead – Tegra System Software

NVIDIA Bristol, United Kingdom

Platform Engineer

Faculty AI London, United Kingdom

Hybrid

Maintenance Technician - McKesson

Ocado United Kingdom

On-site

Principal Software Engineer - Engineering Applications

PhysicsX London, United Kingdom

On-site

Seniority: Lead
Posted: 20 Mar 2026 (Last month)

Save job

Create job alert

Applications closed

Join us at Entrust

At Entrust, we’re shaping the future of identity centric security solutions. From our comprehensive portfolio of solutions to our flexible, global workplace, we empower careers, foster collaboration, and build solutions that help keep the world moving safely.

Get to Know Us

Headquartered in Minnesota, Entrust is an industry leader in identity-centric security solutions, serving over 150 countries with cutting-edge, scalable technologies. But our secret weapon? Our people. It’s the curiosity, dedication, and innovation that drive our success and help us anticipate the future.

About This Role

This is a Product Reliability position, not an infrastructure SRE role. Our DevOps team manages the infrastructure platform; this role focuses on application and service-level reliability, working directly with product engineers.

This is the first role of its kind in product engineering. Reporting to the VP of Product Engineering for Consumer Identity, you’ll drive reliability efforts across the team: defining the roadmap, prioritizing initiatives, and partnering with engineering directors and senior ICs to deliver them.

Why Join Us

Greenfield opportunity: You’ll define Product Reliability as a discipline here. Build the playbook, not inherit one.
High-impact domain: Consumer Identity powers identity verification and biometric authentication for some of the world’s largest financial institutions. Our reliability directly impacts fraud prevention and customer onboarding at scale.
Real authority: Direct line to VP Engineering, budget for tooling, seat at architecture council and service reviews.
Strong foundation: We’re not firefighting. 99.98% uptime means you’re optimizing, not triaging chaos.
Technical depth: Work across ML pipelines, computer vision systems, and mobile SDKs (not just YAML and dashboards).
Ownership culture: Engineers own their services end-to-end; you’ll amplify that, not replace it

Experience Level

Staff SRE

8+ years in software engineering
4+ years in reliability/SRE
Drives reliability initiatives across multiple teams; hands-on with complex systems

Principal SRE

15+ years in software engineering
6+ years in reliability/SRE
Sets technical direction org-wide; influences business-unit-level reliability strategy

We’re open to either level. Scope and compensation will match your experience. Principal candidates should demonstrate cross-org impact and a track record of building reliability programs from scratch.

Current State Incident Analysis (2020–2025)

Postmortem volume peaked in 2023, down 48% since then despite increased release cadence
P0-to-P1+ ratio remains stable despite lower overall incident volume
65% change-induced incidents (deployments, migrations, config changes); 35% organic (third-party outages, expirations, attacks)
Change-induced ratio improved modestly: 69% → 62%
Detection time: 35 min → 18 min
Customer-first detection: 40% → 22%

Availability Targets

2024 & 2025 average uptime: 99.98% (as available in our public status page)
Goal: Consistent 99.99% (four nines) average uptime, SLO breach reductions

System Simplification

We’re reducing system complexity to narrow the reliability target area:
Microservices (K8s deployments/rollouts) reduced 29% from peak, with further cuts planned for 2026
Goal: Smaller footprint, higher reliability, lower cost for new regions

Role Objectives

Primary goal: Improve release safety, reduce releases that cause downtime or SLO degradation.
We already have foundational systems in place:
Automated test coverage and crowd testing
A/B testing and dark canaries
Progressive rollouts (infrastructure and application level)
Back-testing against historical data
To consistently exceed four nines, we need to mature these systems and build new capabilities.

Ideal Candidate Profile

Mindset

Passionate about reliability as a discipline, not just a checkbox
Focused on reliability, not product features, but willing to learn the product to understand impact
Hands-on: eager to build tooling and systems
Pragmatic about balancing reliability with development velocity

Required Skills

Software Engineering

Strong software engineering in at least one of our backend languages (Python, Ruby, Node.js); able to navigate most of our codebase
Experience building reliability tooling: progressive delivery, automated rollbacks, monitoring/alerting

Reliability Patterns

Deep knowledge of resilience patterns: circuit breakers, bulkheads, back-pressure, retries with backoff, rate limiting, load shedding, graceful degradation
Solid incident management and blameless postmortem practices

Observability

Proficiency with observability: distributed tracing, structured logging, metrics instrumentation
Uses data to drive decisions: experienced with SLIs, SLOs, and error budgets

Communication

Skilled at influencing without authority
Able to hold deep technical reliability discussions with senior ICs

Nice-to-Have

Experience with chaos engineering (fault injection, game days, controlled failure experiments)
ML system reliability experience (mixed I/O and CPU-bound workloads, non-deterministic behavior, model serving)
Familiarity with our specific stack (Datadog, Kubernetes, AWS, GitLab CI/CD)
Experience leveraging LLMs for code analysis, design doc review, or automated runbook generation
On-call experience in a high-availability environment

Our Stack

Backend: Python, Ruby on Rails, Node.js
Frontend: React, TypeScript
Mobile: Swift (iOS), Kotlin (Android), React Native
Infrastructure: AWS, Kubernetes, Terraform, SNS, SQS
Databases: PostgreSQL (Aurora), Redis, OpenSearch
Observability: Datadog, Splunk, Sentry
ML: PyTorch, TensorFlow
CI/CD: GitLab (on-prem)

#LI-JS1

At Entrust, we don’t just offer jobs – we offer career journeys. Here is what you can expect when you join our team:

Career Growth: Whether you’re a budding developer or a seasoned expert, we’re invested in your professional journey. With learning-forward initiatives and exciting challenges, your growth is our priority.
Flexibility: Life is all about balance. Whether you’re remote, hybrid, or on-site, we offer flexible options that fit your lifestyle.

Collaboration: Here, your voice matters. Our teams thrive on sharing ideas, brainstorming solutions, and working together to build a better tomorrow.

We believe in securing identities—but it doesn’t stop there. At Entrust, we’re passionate about valuing all identities. Our culture is built on diversity, inclusion, and respect. From unconscious bias training for our leaders to global affinity groups that connect colleagues across the globe, we’re creating a community where everyone is encouraged to be themselves.

Ready to Make an Impact?

If you’re excited by the prospect of innovating, growing your career, and collaborating in a dynamic environment, Entrust is the place for you. Join us in making a difference. Let’s build a more secure world—together.

Apply today!

For more information, visit www.entrust.com. Follow us on, LinkedIn, Facebook, Instagram, and YouTube

For US roles, or where applicable:

Entrust is an EEO/AA/Disabled/Veterans Employer

For Canadian roles, or where applicable:

Entrust values diversity and inclusion and we are committed to building a diverse workforce with wide perspectives and innovative ideas. We welcome applications from qualified individuals of all backgrounds, and we strive to provide an accessible experience for candidates of all abilities.

If you require an accommodation, contact .

Recruiter:

Jack Steib

Industry Insights

Discover insightful articles, industry insights, expert tips, and curated resources.

Apr 9, 2026

Products

Where to Advertise Machine Learning Jobs in the UK (2026 Guide)

Advertising machine learning jobs in the UK requires a different approach to most technical hiring. The candidate pool is small, highly specialised and in demand across AI labs, financial services, healthcare, autonomous systems and consumer technology simultaneously. Machine learning engineers and researchers move between roles through professional networks, conference communities and specialist platforms — not general job boards where ML roles compete with unrelated software engineering positions for the same audience. This guide, published by MachineLearningJobs.co.uk, covers where to advertise machine learning roles in the UK in 2026, how the main platforms compare, what employers should expect to pay, and what the data says about hiring across different role types.

Apr 5, 2026

Jobs

Machine Learning Jobs UK 2026: What to Expect Over the Next 3 Years

Machine learning has undergone a transformation that few technology disciplines can match. In the space of three years it has moved from a specialism sitting at the edges of most organisations' technology strategies to a capability that sits at the centre of them. The tools have changed, the expectations have shifted, and the range of industries treating machine learning as a core business function — rather than an experimental one — has expanded dramatically. For job seekers, this creates both opportunity and complexity in roughly equal measure. The machine learning jobs market of 2026 is significantly larger than it was three years ago, but it is also significantly more demanding. Employers have developed more sophisticated expectations, the technical bar for specialist roles has risen, and the landscape of tools, frameworks, and architectural patterns that practitioners are expected to know has broadened considerably. The candidates who will thrive over the next three years are those who understand where the discipline is heading — which specialisms are attracting the most investment, which technologies are reshaping what machine learning engineers and researchers are expected to build, and how the definition of a machine learning career is evolving beyond the model-building core toward a much wider range of roles across the full ML lifecycle. This article breaks down what the UK machine learning jobs market is likely to look like through to 2028 — covering the titles emerging right now, the technologies driving employer demand, the skills that will matter most, and how to position your career ahead of the curve.

Mar 24, 2026

Jobs

New Machine Learning Employers to Watch in 2026: UK and Global Companies Driving ML Innovation

Machine learning (ML) has transitioned from a specialised field into a core business capability. In 2026, organisations across healthcare, finance, robotics, autonomous systems, natural language processing, and analytics are expanding their machine learning teams to build scalable intelligent products and services. For professionals exploring opportunities on www.MachineLearningJobs.co.uk , understanding the companies that are scaling, winning investment, or securing high‑impact contracts is crucial. This article highlights the new and high‑growth machine learning employers to watch in 2026, focusing on UK innovators, international firms with significant UK presence, and global platforms investing in machine learning talent locally.