Responsibilities
- Own the platform architecture end to end, including architecture proposals, high-level and low-level designs, driven through review and kept current as the baseline.
- Lead and line-manage a distributed team across backend (Node.js), frontend (React), DevOps, QA and documentation, covering engineering standards, code and design review, release gates, one-to-ones, growth and performance input.
- Design, build and operate a managed Slurm service for research users, including controller and accounting, partitions and login nodes, node onboarding and acceptance, driver and CUDA baseline and upgrades, stalled-job and node-health detection, drain and autohealing, storage visibility, identity and isolation.
- Own cluster bootstrap and lifecycle on partner-provided bare metal, NVIDIA GPU Operator and Network Operator, VM-based GPU isolation (KubeVirt and VFIO), and day-2 operations including upgrades, backup and recovery, and node replacement.
- Manage serving architecture for managed inference at scale, including multi-GPU and multi-node parallelism, autoscaling, request routing and endpoint reliability, and confidential-compute-capable capacity for sensitive workloads.
- Implement observability and operations with metrics, logging, alerting and SLOs across control plane, GPU fleet and application tiers, incident response and post-incident review, and an on-call model a small team can sustain.
- Serve as primary technical interface to infrastructure partners and vendors, turning requirements into written specifications and acceptance tests, running escalations to closure, and providing technical input to capacity planning and hardware sourcing.
- Work directly with research, model-training and product teams to translate their workloads into platform requirements, and broker capacity when it is short.
- Complete the platform team and set the technical bar for the engineers who join it.
Requirements
- Eight or more years of hands-on engineering, including at least three leading teams that build and operate infrastructure platforms other teams depend on.
- Bachelor's or Master's degree in computer science or engineering, or equivalent practical experience.
- Hands-on experience with Slurm at scale, including running slurmctld and slurmdbd for real users, managing partitions, QoS and priority, accounting, prolog and epilog, node health scripting, and upgrades with jobs on the system; ideally has operated an HPC or GPU training cluster for a research population.
- Experience with GPU fleet operation on bare metal, including NVIDIA driver and CUDA lifecycle, Fabric Manager and NVSwitch behaviour on SXM systems, DCGM-based health and utilisation, MIG, and node burn-in and acceptance.
- Knowledge of high-performance interconnects, including InfiniBand fabric and subnet configuration, RDMA, SR-IOV, and diagnosing multi-node NCCL performance problems.
- Depth in Linux systems, including kernel modules and drivers, PCIe passthrough and vfio-pci, cgroups and namespaces, and performance tuning for compute-heavy workloads.
- Experience with production Kubernetes operation, not just deployment, including control plane, upgrades, CNI and CSI, operators and custom controllers, and multi-tenancy design.
- Experience with HPC storage and data movement, including shared filesystems (VAST, Lustre, NFS), node-local NVMe caching, and distributing large model weights and datasets across many nodes.
- Experience with observability and operations using Prometheus, Grafana and Loki or equivalents, SLOs, incident response and post-incident review.
- Working fluency in JavaScript and Node.js sufficient to review a control plane, CLI and worker services with authority and to make architecture decisions on them; not a feature-development requirement.
- A shipped platform with real users, such as a multi-tenant IaaS or PaaS, or a research computing service, including resource isolation, quotas, usage metering, and user-facing API and CLI surfaces.
- Leadership that stays in the code, including people management across time zones, cross-track review, written architecture decisions with alternatives recorded, and the ability to tell a partner or an executive no with reasons.
- Excellent written and spoken English; most partner and leadership work happens in writing.
- Fully remote, based between UTC and UTC+5:30 so the working day overlaps both Europe and India, where the team and its partners work; occasional travel to partner sites and team events.
Nice to Have
- Experience with Slurm operators on Kubernetes (Soperator, Slinky) or Kubernetes-native schedulers (Kueue, Volcano, KAI, Kubeflow Trainer).
- Experience with modern serving stacks (vLLM, SGLang, TensorRT-LLM), including parallelism strategies, quantisation trade-offs, and GPU memory planning.
- Experience with VM and container isolation for multi-tenant GPU compute (KubeVirt, Kata Containers, QEMU and KVM, Firecracker) and confidential computing (Intel TDX, AMD SEV-SNP, NVIDIA confidential-compute mode).
- Experience with Cluster API and kubeadm, Cilium, NVSentinel-class autohealing, infrastructure as code and GitOps.
- Time on the operator side of a GPU cloud, a national or university HPC centre, or an AI lab's platform team.
- Background in peer-to-peer or distributed systems.
- Experience with a hardware provider who provisions but does not operate, and turning that relationship into a written contract with acceptance tests.
Work Arrangement
Remote (Worldwide) — Worldwide
Other
- Excellent written and spoken English required.
- Fully remote, based between UTC and UTC+5:30.
- Occasional travel to partner sites and team events.