Skip to main content
Posted 21 July, 2026

Senior Site Reliability Engineer

VIQU IT Recruitment
Milton Keynes, ENG, GB Full Time

Job Description Senior Site Reliability Engineer\nUp to £70,000 plus bonus and on call allowance\nMilton Keynes (2 days on site a week) \nVIQU have...

Job Description

Senior Site Reliability Engineer

\n

Up to £70,000 plus bonus and on call allowance

\n

Milton Keynes (2 days on site a week)

\n

VIQU have partnered with a well-established B2B SaaS company who are going through a significant platform transformation. and so are hiring for a Senior Site Reliability Engineer to build stability, respond to live incidents, and assist with system upkeep. The role will also play a key part in on implementing and adopting new tooling and processes surrounding the wider transformation.

\n

This is a genuine opportunity to own and operate how the cloud function works, and progress into a team lead position as the team grows.

\n

Experience required for the Senior Site Reliability Engineer

\n
    \n
  • Previous experience as a Site Reliability Engineer or similar (cloud, infrastructure, DevOps or platform engineering) within a customer facing environment – e.g SaaS or MSP.
  • \n
  • Strong hands-on experience with both Azure, and on-premise virtual machines.
  • \n
\n
    \n
  • Experience withInfrastructure as Code / Terraform, Container orchestration (Kubernetes or AKS), and Monitoring and observability tooling (Prometheus, Grafana, Datadog, or Azure Monitor).
  • \n
  • Ability to implement new processes, and tools, ensuring the wider development and support teams adopts new ways of working.
  • \n
  • Ability to communicate across internal teams and external customers.
  • \n
  • Skilled in networking across both cloud (Azure) and on premise environments.
  • \n
  • Either Windows or Linux systems administration skills (Linux preferred).
  • \n
  • Previous use of AI tools to enhance efficiency.
  • \n
\n

Job Duties of the Senior Site Reliability Engineer

\n
    \n
  • Utilise various technologies (Terraform, Kubernetes ect) to manage provision, and configure servers and networks, and automate application lifecycles.
  • \n
  • Regularly use Datadog and other observability tools for application performance monitoring.
  • \n
  • Implement new ways of working, helping to shape how the organisation responds and recovers to incidents.
  • \n
  • Take ownership of incident resolutions.
  • \n
\n
    \n
  • Actively drive down key reliability metrics (MTTR, incident frequency, on-call toil) by evaluating key incidents.
  • \n
  • Work on an a on call rota, ensuring you are available to respond to incidents during this time.
  • \n
\n
    \n
  • Identify areas for automation and help implement changes that raise the bar for reliability.
  • \n
\n

\n

Apply now to speak with VIQU IT in confidence. Or reach out to Jack McManus via the

\n

Do you know someone great? We’ll thank you with up to £1,000 if your referral is successful (terms apply). For more exciting roles and opportunities like this, please follow us on LinkedIn @VIQU IT Recruitment

This listing expired on 22 Jul. Applications are no longer accepted.

Below are some other jobs we think you might be interested in.