Responsibilities
- Own the end-to-end platform architecture, including architecture proposals, high-level and low-level designs, driven through review and kept current as the baseline.
- Lead and line-manage a distributed team across backend (Node.js), frontend (React), DevOps, QA, and documentation, setting engineering standards, code and design review, release gates, one-to-ones, and growth and performance input.
- Design, build, and operate a managed Slurm service for research users, covering controller and accounting, partitions and login nodes, node onboarding and acceptance, driver and CUDA baseline and upgrades, stalled-job and node-health detection, drain and autohealing, storage visibility, identity, and isolation.
- Own cluster bootstrap and lifecycle on partner-provided bare metal, including NVIDIA GPU Operator and Network Operator, VM-based GPU isolation (KubeVirt and VFIO), and day-2 operations such as upgrades, backup and recovery, and node replacement.
- Manage serving architecture for inference at scale, including multi-GPU and multi-node parallelism, autoscaling, request routing, endpoint reliability, and confidential-compute-capable capacity for sensitive workloads.
- Handle observability and operations, including metrics, logging, alerting, and SLOs across control plane, GPU fleet, and application tiers, incident response and post-incident review, and an on-call model a small team can sustain.
- Serve as the primary technical interface to infrastructure partners and vendors, turning requirements into written specifications and acceptance tests, running escalations to closure, and providing technical input to capacity planning and hardware sourcing.
- Work directly with research, model-training, and product teams to translate their workloads into platform requirements and broker capacity when it is short.
- Complete the platform team and set the technical bar for the engineers who join it.
Compensation
Not specified
Work Arrangement
Remote (Worldwide) — Remote (UTC to UTC+5:30)
Team
Fully remote, based between UTC and UTC+5:30 so the working day overlaps both Europe and India, where the team and its partners work. Occasional travel to partner sites and team events.
Other
- Fully remote, based between UTC and UTC+5:30 so the working day overlaps both Europe and India, where the team and its partners work.
- Occasional travel to partner sites and team events.
Not specified