Responsibilities
- Ensure consistent performance and stability of the live production environment by managing EKS, ArgoCD, and AWS components, optimizing autoscaling behavior, and refining incident response to prevent recurring issues.
- Implement reliable progressive delivery workflows using Argo Rollouts, enabling rapid rollback within seconds or minutes when deployments fail.
- Develop internal developer tools as polished products, improving the local development setup, per-PR staging environments, and CLI utilities to support fast, concurrent workflows including AI-assisted development.
- Reduce repetitive operational tasks by applying AI to automate low-risk processes, allowing engineers to focus on complex challenges while keeping humans in control of infrastructure changes.
- Maintain up-to-date infrastructure by leading version upgrades for Kubernetes, EKS, Helm, Terraform-managed resources, and application runtimes, while managing security vulnerabilities on the infrastructure side.
- Enhance cost efficiency by improving visibility into cloud usage through Datadog monitoring, rightsizing AWS resources, and refining log filtering to ensure observability scales efficiently.
- Collaborate with engineering teams building on the platform by improving documentation standards, participating in shared security responsibilities, and advancing team-wide ownership to reduce knowledge silos.
Work Arrangement
Remote (Worldwide)
Work Arrangement
- Fully remote position with global availability; candidates from any location are welcome provided they have at least four hours of daily overlap with EMEA time zones.
- Occasional in-person gatherings occur once or twice annually and are strongly encouraged to strengthen team relationships within a small, distributed group.
- On-call duties are shared across all engineering staff, with each engineer taking a week-long rotation approximately once every quarter.