Responsibilities
- Own the cloud foundation. Design, build, and own our AWS architecture and the Terraform/IaC that defines it — accounts, networking, identity, core services — as reproducible, reviewed, version-controlled infrastructure.
- Build the self-service platform. Create the golden paths and internal developer-platform tooling that let engineers and researchers provision, deploy, and ship with minimal friction and sensible guardrails.
- Own CI/CD and the deployment fabric. Build and maintain the pipelines and deployment substrate — including the path that gets our models and SDK into cloud and on-prem/customer environments reliably.
- Engineer security and networking into the platform. Make security and a sound network a property of the paved road, not a checklist: IaC for access management and access control, secrets management, least-privilege defaults, network segmentation and private connectivity — so doing the right thing is the easy thing.
- Set the standards. Create, influence, and review the architecture, standards, and processes the platform team works to; decide, with a product lens, what becomes platform versus what stays an experiment.
- Make it observable and cost-aware. Build the metrics, alerting, and cost visibility that keep the platform reliable and economical as scale grows.
- Support research at scale. Bring platform know-how to research workflows and help them self-serve — supporting scale, not owning research-infra execution.
Requirements
- Proficiency in AWS — hands-on production experience designing and running cloud infrastructure for real workloads.
- Proficiency in Infrastructure-as-Code — Terraform in production (configuration-as-code such as Ansible a plus); you treat infrastructure as reviewed, reproducible software.
- Proficiency in Kubernetes and containers — hands-on with container orchestration and the components around it.
- Strong CI/CD and developer-experience instincts — GitHub Actions, ArgoCD/Flux or equivalent; you've built pipelines and internal tooling that other engineers depend on.
- Security-engineering capability — you implement access control, secrets, and least-privilege in code, and reason about the security posture of what you build.
- Strong networking — hands-on with cloud and hybrid networking (VPC design, segmentation, private connectivity, DNS, load balancing, ingress/egress control). You treat the network as a first-class security and reliability surface — this complements the security-engineering work directly.
- A platform-as-a-product mindset — you can tell what deserves to be a productized paved road versus a script, ship opinionated defaults, and enable teams rather than gate them.
- Solid software engineering — Python, Go, or a comparable language, with sound engineering practices.
Nice to Have
- Experience with cloud GPU compute for ML training; exposure to HPC (SLURM, InfiniBand, bare-metal GPU) — useful for understanding research workflows even though we run cloud-first.
- Modern observability stacks (Grafana, Prometheus/VictoriaMetrics, Loki).
- Experience in pharma, biotech, healthcare, or another regulated-data environment.
- Exposure to SOC 2, ISO 27001, or GDPR controls (we build the engineering so certification is easy later).
- Experience contributing to or maintaining open-source projects.
- Start-up or fast-scaling environment, with a high degree of ownership.
Additional Information
- CV must be submitted in English
- Remote-friendly working
- Flexible working