Site Reliability Engineer
Keep our global cloud platform reliable, scalable, and resilient.
Do you enjoy solving complex production challenges before they turn into incidents? Are you the kind of engineer who automates repetitive tasks, improves reliability through engineering, and believes that every outage is an opportunity to build a better system? And are you an experienced, hands-on engineer who wants to stay close to the technology rather than move into people management?
We're looking for a Senior Site Reliability Engineer to join the Global Platform Team at Vertage. In this role, you'll help build and operate the cloud platform that powers our global Workforce-as-a-Service (WaaS) ecosystem. You'll work alongside Cloud Engineers, Platform Engineers, Data Engineers, and AI specialists to improve platform resilience, reduce operational overhead, and ensure that our Azure-based engineering environment remains reliable, scalable, and resilient. This is a hands-on individual contributor role, with no people management responsibilities.
Your impact
As a Senior Site Reliability Engineer, your focus is to keep our platforms healthy, reliable, and easy to operate while continuously improving how we build, run and recover production systems.
You'll help build the engineering foundations behind our Headless Data Architecture (HDA), which runs on Azure and Databricks, as well as the Custom Apps Infrastructure (CA) that powers integrations, internal applications, and operational workflows across our international organization.
Instead of spending your days reacting to incidents, you'll focus on preventing them through automation, observability, and reliability engineering. You'll reduce operational burden, improve platform resilience, and build systems that scale, recover automatically whenever possible, and provide engineering teams across Cloud, Data, and AI with the visibility they need to run production workloads with confidence.
What you will do
Improve the reliability, availability, and performance of our Azure platform and production environments.
Build and improve monitoring, logging, and alerting using Grafana, OpenTelemetry, Azure Monitor, and Log Analytics.
Automate operational tasks and eliminate repetitive manual work using Infrastructure as Code and scripting.
Design self-healing capabilities and automated remediation to reduce incidents and improve recovery times.
Investigate production incidents, perform root cause analyses, and implement long-term improvements.
Define, measure, and improve Service Level Indicators (SLIs) and Service Level Objectives (SLOs).
Optimize platform performance, scalability, and operational efficiency.
Work closely with Cloud Engineers to improve platform design, resilience, security, and operational reliability.
Support Data and AI teams by improving the reliability of Azure Databricks environments.
Help shape and raise engineering standards around observability, automation, and operational excellence.
Continuously look for opportunities to reduce operational complexity and improve developer productivity.
Act as a technical point of reference for reliability engineering, helping shape technical decisions while remaining hands-on in the technology.
About the role
As part of the Global Platform Team, you'll work alongside engineers in Cloud, Data, and AI to improve the reliability of our Azure-based platform. Using technologies such as Kubernetes, Terraform, Databricks, GitHub Actions, Grafana, and OpenTelemetry, you'll help ensure that our global Workforce-as-a-Service ecosystem remains reliable, scalable, and resilient.
This is a hands-on individual contributor role. You will have significant technical ownership and will be expected to get into the detail when needed, from troubleshooting complex production issues to improving automation, observability and system reliability. You will act as a technical reference point for others, but you will not have people management responsibilities.
About Vertage
Vertage is one of Europe's leading providers of workforce and talent solutions. Operating across multiple countries, we are transforming into a cloud-native, AI-powered organization that connects people, technology, and data through a modern digital platform. The Global Platform Team is at the heart of that transformation, enabling engineering teams across Cloud, Data, AI, and Software Engineering to build and operate scalable solutions for the future.
You're passionate about building reliable systems and solving operational challenges through engineering rather than manual intervention. You enjoy understanding how distributed systems behave, thrive in cloud-native environments, and are always looking for ways to improve automation, resilience, and observability. You stay calm under pressure, take ownership of problems, and enjoy collaborating with others to continuously improve the platform's reliability.
You have significant experience operating production systems and still enjoy getting hands-on. You are comfortable working independently, making technical decisions and going deep into complex problems. While you may naturally act as a technical reference point for others, you are not looking for a people management role.
7+ years of experience as a Site Reliability Engineer, Platform Engineer, DevOps Engineer, or Cloud Engineer, with significant hands-on experience in production environments.
Strong hands-on experience with Microsoft Azure.
Experience with Infrastructure as Code using Terraform.
Experience building and maintaining CI/CD pipelines using GitHub Actions or Azure DevOps.
Experience with observability tools such as Grafana, OpenTelemetry, Azure Monitor, or Log Analytics.
Strong scripting skills using Python, Bash, or similar languages.
Experience supporting distributed cloud platforms in production.
Experience with incident management, root cause analysis, and post-incident improvements.
Familiarity with GitOps principles and modern deployment practices.
Experience with Azure Databricks is a strong advantage.
Experience with SnapLogic or similar integration platforms is a plus.
A strong preference for staying hands-on and solving technical problems directly, rather than moving into people management.
Practical information
Location: Luton, Bedfordshire, United Kingdom
Working model: Hybrid, 60/40, with regular presence at our Luton office.
Language: English is required.
Work authorisation: You must already have the right to work in the UK. We are unable to provide visa sponsorship for this position.
Interested?
If you're an experienced SRE who enjoys solving complex engineering problems, staying hands-on with technology, and taking ownership of production reliability, we'd love to hear from you.
Mobility budget or Lease car: No
Brand: Vertage
Recommended Jobs
Group Workout Instructor
We are Places for People Group, we're a social enterprise that believes it's people that make a community. That's why we build homes and deliver services for everyone in the community to thrive. At P…
Head of Office
London based Family Office is seeking a Head of Office to manage the financial affairs of two Founders, and their related UK, US and Caribbean entities. Reporting directly to the Founders, the HOO wi…
Early Years Apprentice - Little Forest Folk Chiswick
Summary Are you a passionate early years practitioner looking to continue to develop your career in an fully outdoor forest nursery setting? Little Forest Folk Chiswick have a great opportunity to…
Construction Litigation Solicitor
Construction Litigation Solicitor/Senior Associate or Legal Director | London | 7+ PQE This leading national law firm is seeking an experienced Construction Litigation Solicitor to join its high…
Team Leader
Incipio curates beautiful spaces with vibrant atmospheres for great times. Juno is a contemporary Italian restaurant and bar launching at Olympia, a space that blends the ease of neighbourhood din…
Marketing Director
Marketing Director – Events GBP60,000 – GBP75,000 Base + Bonus Hybrid – London The Company Global events business seeks highly accomplished Marketing Director to lead a marketing team and…
Chief Actuary (Part-Time)
We've partnered with a leading global insurance provider to find a Part-Time Chief Actuary to join their UK operation. This is a rare opportunity for a senior actuarial leader to take on a high-impac…
Social Worker - Neighbourhood Team
Liquid Personnel is looking for a skilled and compassionate Senior Social Worker to join its client's Mental Health Team based in London Borough of Lewisham. This role involves leading on complex …
SEN Co-ordinator
Our client Enfield Council is looking for a SEN Co-ordinator to join their team. ~ Full Time working from home (Hybrid working possible in the Civic Centre) ~£278.32 per day - rate to be confi…
Senior Legal Counsel
NBCUniversal is one of the world's leading media and entertainment companies. We create world-class content, which we distribute across our portfolio of film, television, and streaming, and bring…