Site Reliability Engineer (Mid / Senior)
Job Description
Site Reliability Engineer (Mid / Senior)
\nSouth West London (Hybrid – 1–2 days onsite)Salary: Competitive + Benefits
\nWe are looking for a Site Reliability Engineer to join a well-established small infrastructure team supporting a highly available, production environment. This is an exciting opportunity to work across a modern, self-hosted platform spanning Kubernetes, physical infrastructure and automation, with a strong focus on Ubuntu-based systems.
\nThe Role
\nAs an SRE, you will play a key role in ensuring the availability, performance, security and resilience of production systems. Working in a small, collaborative team, you’ll take ownership of day-to-day platform operations, incident response and continuous improvement, while partnering closely with development teams to deliver reliable and scalable services.
\nKey Responsibilities
\n- \n
- Administer and maintain Linux (Ubuntu) server environments \n
- Manage self-hosted Kubernetes clusters and supporting infrastructure \n
- Support on-premise infrastructure including physical servers and virtualisation platforms \n
- Administer storage solutions including NFS, iSCSI and object storage \n
- Build and maintain automation using Ansible or similar IaC tools \n
- Develop operational tooling using Bash and Python \n
- Monitor system health using tools such as Prometheus, Grafana, Zabbix or Nagios \n
- Investigate and resolve production incidents (on-call rota involved) \n
- Implement security hardening and infrastructure best practices \n
- Manage backup and disaster recovery processes and regular testing \n
- Support and improve CI/CD pipelines and deployment processes \n
- Collaborate with engineering teams to improve reliability and performance \n
Essential Skills
\n- \n
- Strong Linux systems administration (Ubuntu preferred) \n
- Experience running production Kubernetes environments \n
- Solid understanding of networking (TCP/IP, DNS, routing, firewalls) \n
- Experience with physical servers and virtualisation platforms \n
- Hands-on experience with Ansible or other IaC tools \n
- Scripting skills in Bash and Python \n
- Experience with monitoring and alerting platforms \n
- Knowledge of Linux storage technologies (NFS, iSCSI) \n
- Experience with backup & disaster recovery \n
- Exposure to Active Directory / Entra ID / endpoint management \n
- Strong troubleshooting and problem-solving skills \n
Desirable Experience
\n- \n
- Object storage, MariaDB or database administration \n
- CI/CD tools such as Jenkins \n
- AWS (S3, Lambda, CloudFront) exposure \n
- Terraform or additional IaC tooling \n
- Experience with Harvester or similar platforms \n
- Knowledge of security, compliance or GDPR \n
Why Apply?
\n- \n
- Work on complex, real-world infrastructure (not just cloud-native) \n
- High ownership in a small, collaborative team \n
- Exposure to a broad modern tech stack across infra, Kubernetes and automation \n
- Hybrid working with a competitive salary package \n
