We’re looking for a Principal Platform Engineer to architect and lead the infrastructure strategy for our next-generation Production ML platform on Google Cloud. In this role, you will be the backbone of our high-performance machine learning workloads, ensuring our systems are elastic, secure, and resilient. You won’t just maintain the status quo; you’ll build the "paved road" for our engineers, automating everything from model deployment to complex networking perimeters. We are a high-trust, outcome-focused team that moves quickly to solve some of the most challenging problems in the ML space.
Core Responsibilities:
- Infrastructure Management: Design, deploy, and maintain elastic scaling cloud infrastructure (GCP) and containerization tools like Kubernetes for high-performance ML workloads.
- CI/CD Pipeline Development and maintenance: Build automated pipelines for training, testing, and deploying machine learning models using tools like Jenkins, GitHub Actions, or Airflow.
- Model Monitoring & Maintenance: Implement observability tools to track model drift, accuracy, latency, and performance degradation in production.
- Collaboration: Bridge the gap between data engineers, ML engineers, Backend and Frontend engineers to ensure smooth production operation.
- ML Observability: Implement comprehensive monitoring for system health (latency/uptime) alongside ML-specific metrics, such as feature drift, prediction accuracy, and data distribution shifts, to ensure long-term model reliability. Non ML workload and production metrics monitoring.
- Deploy tools that empower individual teams to monitor their workloads.
- Participate in on-call rotation, help manage posture to ensure compliance with standards such as SOC.
What you bring to the table:
- Senior Expertise: 8 - 10+ years in DevOps/Platform Engineering, with at least 2 years of experience specifically operating and maintaining production ML workloads.
- GCP & K8s Mastery: Deep, hands-on experience with GCP (VPC-SC, IAM, Organization Policies) and GKE (Cluster topology, Helm, Kustomize, and in-cluster operators like ArgoCD).
- Service Mesh Excellence: High proficiency with Istio (VirtualServices, mTLS, sidecar injection) and API Gateways (specifically Kong).
- Infrastructure as Code: Expert-level Terraform skills, specifically using an Atlantis/GitOps workflow across a massive, multi-hundred-file estate.
- Secrets & Identity: Experience managing enterprise-grade identity and secrets (Auth0, Dex, ESO, or SOPS).
- Data/ML Tooling: Experience operating Airflow in production and an ML-serving stack (e.g., Triton, vLLM, MLflow).
- Database Management: Comfortable managing Cloud SQL (PostgreSQL), BigQuery, and in-cluster datastores like Elasticsearch or ClickHouse.
- At least an upper-intermediate level of spoken and written English.
It would be great if you also had:
- ML Observability: Past experience with continuous monitoring of model accuracy and detecting data/concept drift.
- Automation Savvy: Experience with Ansible for cluster bootstrap and recovery.
- Advanced Certifications: Kubernetes (CKA/CKS) or GCP Professional Cloud Architect/Security Engineer certifications.
- Modern Stack Exposure: Familiarity with Loki, Grafana, or managing ClickHouse at scale.
Apply on company website Remote Estonia Remote (Country) Full-time