Responsibilities
- Design, deploy, and maintain elastic scaling cloud infrastructure (GCP) and containerization tools like Kubernetes for high-performance ML workloads.
- Build automated pipelines for training, testing, and deploying machine learning models using tools like Jenkins, GitHub Actions, or Airflow.
- Implement observability tools to track model drift, accuracy, latency, and performance degradation in production.
- Bridge the gap between data engineers, ML engineers, Backend and Frontend engineers to ensure smooth production operation.
- Implement comprehensive monitoring for system health (latency/uptime) alongside ML-specific metrics, such as feature drift, prediction accuracy, and data distribution shifts, to ensure long-term model reliability.
- Deploy tools that empower individual teams to monitor their workloads.
- Participate in on-call rotation, help manage posture to ensure compliance with standards such as SOC.
Requirements
- Senior Expertise: Operating and maintaining production ML workloads, and DevOps/Platform Engineering.
- Data/ML Tooling: Experience operating Airflow in production and an ML-serving stack (e.g., Triton, vLLM, MLflow).
- GCP & K8s Mastery: Deep, hands-on experience with GCP (VPC-SC, IAM, Organization Policies) and GKE (Cluster topology, Helm, Kustomize, and in-cluster operators like ArgoCD).
- Service Mesh Excellence: High proficiency with Istio (VirtualServices, mTLS, sidecar injection) and API Gateways (specifically Kong).
- Infrastructure as Code: Expert-level Terraform skills, specifically using an Atlantis/GitOps workflow across a massive, multi-hundred-file estate.
- Secrets & Identity: Experience managing enterprise-grade identity and secrets (Auth0, Dex, ESO, or SOPS).
- Database Management: Comfortable managing Cloud SQL (PostgreSQL), BigQuery, and in-cluster datastores like Elasticsearch or ClickHouse.
- At least an upper-intermediate level of spoken and written English.
Nice to Have
- Automation Savvy: Experience with Ansible for cluster bootstrap and recovery.
- Advanced Certifications: Kubernetes (CKA/CKS) or GCP Professional Cloud Architect/Security Engineer certifications.
- Modern Stack Exposure: Familiarity with Loki, Grafana, or managing ClickHouse at scale.
Benefits
- Solve real customer problems with immediate cyber protection solutions.
- See your impact daily in a scrappy, nimble organization where individual contributions are valued.
- Accelerate your career by learning new technologies, products, and markets in a fast-paced, growth-oriented environment.
- Work with other talented people at a company where people matter.
- Opportunity to put your fingerprint on an organization and leapfrog your growth.
Additional Information
- No employee or applicant will face discrimination or harassment based on race, color, ancestry, national origin, religion, age, gender, marital domestic partner status, sexual orientation, gender identity, disability status, or veteran status.
- Point Wild is committed to being an inclusive community where all feel welcome.