Skip to main content
Posted 02 August, 2026

Senior HPC & Cloud Engineer - AI/ML Compute Infrastructure

Sentinel
Oxford, ENG, GB Full Time

Job Description

\n

Senior HPC & Cloud Engineer - AI/ML Compute Infrastructure
\n Oxford | Hybrid (3 days in office) | Competitive, DOE

\n

Sentinel is recruiting for several senior/staff-level engineers to design, build and operate a hybrid GPU compute environment, combining on-prem HPC clusters with public cloud infrastructure for large-scale AI research workloads.

\n

Responsibilities:

\n
    \n
  • Build and operate high-performance GPU training/inference clusters, including scheduling, isolation and automated life cycle management
  • \n
  • Design high-throughput data paths across compute and storage, including parallel filesystems (eg Lustre)
  • \n
  • Benchmark and resolve performance bottlenecks across compute, network and orchestration layers
  • \n
  • Implement observability, resilience and security controls for a compliance-conscious research environment
  • \n
  • Work with research and applied ML teams to forecast GPU/storage capacity and streamline experimentation pipelines
  • \n
\n

Requirements:

\n
    \n
  • Experience with HPC/GPU clusters, including a strong understanding of GPU architecture, high-speed networking and distributed training performance
  • \n
  • Cloud platform experience (Azure, GCP, AWS or other)
  • \n
  • Experience with containerisation/Kubernetes
  • \n
  • Working knowledge of IaC and CI/CD (eg Terraform, Argo CD)
  • \n