Infrastructure Operations
• Manage and maintain cloud infrastructure on GCP (primary), with exposure to AWS
• Perform day-to-day Kubernetes (GKS) operations pod troubleshooting, node management, scaling, and resource tuning
• Execute compliance infrastructure tasks instance provisioning, security group updates, DNS changes, and certificate renewals
CI/CD & Deployments
• Support and maintain CI/CD pipelines built on ArgoCD and GitHub Actions
• Assist engineering teams with onboarding applications to the CI/CD platform
• Troubleshoot build failures, deployment issues, and rollback scenarios
• Ensure deployment hygiene proper tagging, versioning, and environment promotion
Monitoring & Incident Response
• Monitor infrastructure health using Prometheus, Grafana, and AlertForge
• Respond to alerts, triage production issues, and execute runbooks
• Perform initial RCA (Root Cause Analysis) for production incidents and escalate when needed
• Maintain and update runbooks and operational documentation
Automation & Improvement
• Write and maintain Terraform modules for infrastructure provisioning
• Automate repetitive operational tasks using shell scripts, Python, or Go
• Identify and implement improvements to reduce toil and manual effort
• Contribute to internal tools and utilities that improve developer experience
Cost & Efficiency
• Own infrastructure cost as a first-class metric drive continuous cost optimisation across compute, network, and storage
Apply on company website Remote, Germany Hybrid Full-time