Posted 02 August, 2026
Senior HPC & Cloud Engineer - AI/ML Compute Infrastructure
Sentinel
Oxford, ENG, GB
Full Time
Job Description
\n
Senior HPC & Cloud Engineer - AI/ML Compute Infrastructure
\n Oxford | Hybrid (3 days in office) | Competitive, DOE
Sentinel is recruiting for several senior/staff-level engineers to design, build and operate a hybrid GPU compute environment, combining on-prem HPC clusters with public cloud infrastructure for large-scale AI research workloads.
\nResponsibilities:
\n- \n
- Build and operate high-performance GPU training/inference clusters, including scheduling, isolation and automated life cycle management \n
- Design high-throughput data paths across compute and storage, including parallel filesystems (eg Lustre) \n
- Benchmark and resolve performance bottlenecks across compute, network and orchestration layers \n
- Implement observability, resilience and security controls for a compliance-conscious research environment \n
- Work with research and applied ML teams to forecast GPU/storage capacity and streamline experimentation pipelines \n
Requirements:
\n- \n
- Experience with HPC/GPU clusters, including a strong understanding of GPU architecture, high-speed networking and distributed training performance \n
- Cloud platform experience (Azure, GCP, AWS or other) \n
- Experience with containerisation/Kubernetes \n
- Working knowledge of IaC and CI/CD (eg Terraform, Argo CD) \n
