Posted 23 July, 2026
site reliability engineer
Enfint
Greater London, ENG, GB
Full Time
Job Description
\n Описание\n
#J-18808-LjbffrEPAM is a global provider of digital engineering, cloud, and AI-enabled transformation services, focusing on complex software product development and digital platform engineering.
Задачи\n- \n
- Lead and scale a global SRE organization, focusing on engineering excellence and team empowerment \n
- Collaborate with product, platform, operations, and security teams to embed reliability within SDLC practices \n
- Define and monitor KPIs for system reliability, performance, and operational efficiency \n
- Advance automation, Infrastructure as Code approaches, and promote self-healing systems using AI/ML techniques \n
- Develop robust incident management frameworks and lead major incident response activities for critical systems \n
- Implement blameless postmortems and deliver systemic improvements across production environments \n
- Establish observability strategies with standardized tooling for metrics, logs, and tracing to support distributed systems \n
- Adopt and enforce SRE practices, including SLIs, SLOs, SLAs, and error budgets across services \n
- Drive resilience strategies with highly available architectures and disaster recovery readiness \n
- Champion an automation-first culture, leveraging CI/CD pipelines and operational tooling to reduce manual processes \n
- \n
- Strong background in Site Reliability Engineering, DevOps, or platform operations in complex, distributed environments \n
- Expertise in observability platforms, troubleshooting distributed systems, and telemetry-driven insights \n
- Hands‑on experience with automation, Infrastructure as Code (Terraform or CloudFormation), and CI/CD practices \n
- Deep understanding of incident management processes, ITSM standards, and ITIL principles \n
- Knowledge of resilience design patterns, high availability, and fault‑tolerant architectures \n
- Familiarity with AI/ML‑driven approaches for operational efficiency and system reliability \n
- Ability to lead transformation, influence across teams, and foster continuous improvement in culture \n
- Nice to have: Experience in financial services or other highly regulated, mission‑critical environments, Certifications in cloud technologies such as AWS, Exposure to AIOps platforms or advanced observability tooling \n
