Responsibilities
- Set cloud strategy and architecture in partnership with the Director of Cloud & Developer Operations and Chief Architect.
- Define reference architectures, standards, and reusable patterns; guide architectural decisions and hold coherent technical direction across teams.
- Apply Cloud Adoption Framework (CAF) and Well-Architected Framework (WAF) thinking; design for scalability, including the SQL data estate beneath the platform.
- Take end-to-end technical ownership of infrastructure transformation projects — cloud modernisation, efficiency, and FinOps initiatives.
- Drive strategic cloud initiatives and major transformation efforts, and push the boundaries of scale, resilience, and efficiency.
- Bring AI-native operations into the lane — agents, bounded automation, and evals, with observability, approvals, containment, and rollback — so reliability and efficiency scale faster than headcount.
- Mentor engineers on architecture, best practices, and complex problem-solving.
- Lead critical technical discussions and champion innovation; strengthen cross-team collaboration.
- Define and publish standards — make implicit expectations explicit and inspectable.
- Partner with business stakeholders so cloud solutions directly support TOPdesk's strategic objectives.
- Refine operational processes so we stay both agile and highly reliable, and support 24/7 continuity of service (including a shared stand-by rotation).
- Identify dependencies early and cooperate with other engineering and non-cloud teams toward shared goals.
Requirements
- Senior technical authority for cloud engineering behind large-scale SaaS platform.
- Hands-on in the estate, setting the direction other engineers build on.
- Form and execute the cloud technical strategy with the Director of Cloud & Developer Operations.
- Lead the hardest cross-team cloud initiatives and infrastructure transformation.
- Bring AI-native ways of working into how the platform is built, operated, and remediated — agents and bounded automation with observability, approvals, containment, and rollback.
- Apply Cloud Adoption Framework (CAF) and Well-Architected Framework (WAF) thinking.
- Design for scalability, including the SQL data estate beneath the platform.
- Take end-to-end technical ownership of infrastructure transformation projects.
- Drive strategic cloud initiatives and major transformation efforts.
- Mentor engineers on architecture, best practices, and complex problem-solving.
- Lead critical technical discussions and champion innovation.
- Define and publish standards — make implicit expectations explicit and inspectable.
- Partner with business stakeholders so cloud solutions directly support strategic objectives.
- Refine operational processes to maintain agility and high reliability, including participation in a shared stand-by rotation.
- Identify dependencies and cooperate with other engineering and non-cloud teams toward shared goals.
Work Arrangement
Remote (City/Region) — Delft
Team
Structure: Teams of social technicians building a robust, scalable cloud platform, working to agile principles with a collaboratively managed backlog and two-week sprints.
Success in your first year
- A major infrastructure transformation is under your technical ownership and delivering — measurably better scale, resilience, or efficiency against a documented baseline.
- Architectural direction for the cloud platform is coherent across teams, with published standards engineers build on.
- Engineers you mentor are visibly stronger; you are consulted during design, not only after incidents.
- AI-native operations — agents, bounded automation, evals, with approvals and rollback — are part of how the lane runs, not a side experiment.
Technical environment
- Claude Code is our current primary agentic engineering environment: agents may plan and execute within defined constraints and governed environments, while engineers stay accountable for problem framing, architecture, acceptance criteria, security, verification, and production outcomes.
- Agent-generated work is backed by automated tests, runtime checks, reviewable evidence, and risk-based human approval — we match deterministic automation, structured workflows, or agents to the nature and risk of the task rather than maximising autonomy for its own sake.
- Cloud: Microsoft Azure (networking, compute, storage) at multi-tenant SaaS scale; Kubernetes / Azure AKS.
- Infrastructure as code: Terraform via CI/CD; configuration management with Puppet and Ansible.
- Observability: metrics and monitoring stacks (e.g. Influx, VictoriaMetrics, Grafana, Nagios).
- Automation: Python and automation tooling — and we expect you to take the stack to the next level, not just operate today's.
- Legacy: Java, MS SQL, heritage architecture — being decomposed.
- Mechanisms: subagents, MCP tool integrations, multi-agent workflows, and shared prompt, agent, and eval libraries.
Additional Information
- Shared stand-by rotation for 24/7 continuity of service.
- Agile principles with collaboratively managed backlog and two-week sprints.