Responsibilities
- Manage and support Linux-based systems using Debian or Ubuntu distributions.
- Deploy, operate, and scale Kubernetes clusters across bare-metal, virtualized, and on-premises environments.
- Handle end-to-end cluster lifecycle management including upgrades, node configuration, networking, storage, and security enhancements.
- Automate infrastructure provisioning and operational tasks using Ansible, Bash, Python, and GitOps methodologies.
- Design and maintain network infrastructure involving VLANs, L2/L3 routing, VPNs, and multi-site connectivity.
- Develop automated bare-metal provisioning workflows using PXE, Preseed, and cloud-init.
- Implement and maintain observability solutions including Prometheus, Grafana, Loki, ELK, and Graylog.
- Lead incident response efforts and manage escalation procedures during platform outages.
- Enhance system uptime and reduce latency across infrastructure layers.
- Define and enforce service level objectives and indicators across physical, virtual, and service layers.
- Refine alerting and monitoring systems to deliver meaningful, actionable insights.
- Organize and maintain on-call schedules ensuring global coverage across time zones.
- Create and update standard operating procedures for consistent operations and maintenance.
- Coordinate physical infrastructure maintenance for Policloud environments, including hardware troubleshooting and data center operations.
- Administer virtualization and orchestration platforms such as OpenStack, Proxmox, and VMware.
- Contribute to the design and evolution of system architecture across all product lines.
- Forecast resource needs based on projected demand and growth trends.
- Partner with development teams to improve code quality and optimize infrastructure usage.
- Collaborate with cross-functional teams including Hivenet, Policloud, and Customer Success.
Requirements
- Extensive hands-on experience running Kubernetes in production settings.
- Advanced networking skills including VLANs, L2/L3 routing, VPNs, and multi-site connectivity—critical for this role.
- Proficient in Linux system administration, particularly on Debian or Ubuntu platforms.
- Solid grasp of networking fundamentals and experience designing complex network topologies.
- Proven experience creating and managing automation pipelines using Ansible, Bash, Python, and Git-based workflows.
- Familiarity with observability tools such as Prometheus, Grafana, ELK, Loki, or Graylog.
- Experience working with virtualization technologies like OpenStack, Proxmox, or VMware.
- Hands-on experience provisioning bare-metal servers using MAAS (Metal as a Service).
- Strong understanding of distributed systems and container orchestration principles.
- Process-driven mindset with ability to create SOPs and operational frameworks independently.
- Background in incident management, escalation protocols, and on-call rotations.
- Capable of working independently in a fast-moving, engineering-focused environment.
- Technical proficiency combined with commitment to team collaboration and values.
Nice to Have
- Experience with service mesh technologies such as Istio or Linkerd, or advanced CNI configurations.
- Knowledge of Cloudflare APIs, DNS automation, or tunnel setup.
- Experience managing GPU-enabled infrastructure, node setup, and scheduling.
- Understanding of security best practices including RBAC, firewalls, and network policies.
- Familiarity with IT asset or license tracking systems.
- Experience collaborating across distributed teams in multiple time zones.
- Background implementing SRE practices and reliability frameworks in scaling organizations.
Benefits
- Fully remote position with flexible working hours
- High-impact role offering autonomy and ownership
- Collaborative, international engineering team
- Access to a modern tech stack focused on reliability and automation
Work Arrangement
Remote (Worldwide)
Team
Collaborative and international engineering team
Other
- Fluent English is required
- Start date: As soon as possible
- Location: Fully remote within EU timezone (CET ±2h)