Santa Clara, Canada On-site Full-time USD 180,000 – 240,000 / year

Gatik AI is hiring a Senior Cloud Infrastructure Engineer

Responsibilities

  • Architect and maintain mission-critical Kubernetes clusters optimized for heavy GPU/TPU workloads.
  • Implement and optimize Kubernetes-native GPU scheduling (NVIDIA GPU Operator) to ensure maximum hardware utilization.
  • Drive the "Everything as Code" philosophy using Terraform, Helm, and cloud-native tools.
  • Deploy Autonomous AI Agents (LangGraph, CrewAI) to monitor cluster health and enable automated triage of hardware failures and NCCL timeouts.
  • Build large-scale pipelines using Apache Airflow, Kafka, and Spark to process raw sensor data into training-ready formats.
  • Implement robust GitOps workflows using ArgoCD, Gitlab CI/CD to automate the deployment of both infrastructure and model artifacts.
  • Maintain deep visibility into infrastructure health and model serving performance using Prometheus, Grafana, and OpenTelemetry.
  • Develop agent-driven workflows to optimize the developer experience, such as automated PR reviewers for Terraform and AI agents that proactively suggest Kubernetes resource-limit adjustments based on model training telemetry.
  • Design and maintain MLFlow and feature store integrations to provide a robust system of record for every model iteration.
  • Build complex, automated model lifecycles using Airflow and Kubernetes to streamline the transition from training to simulation.
  • Support the deployment of models into simulation and production environments using Triton Inference Server, Ray Serve, and ONNX Runtime.
  • Enable researchers to scale models (VLA, World Models) across multi-node setups using PyTorch Distributed (TorchElastic), Ray Train, and Horovod.
  • Optimize low-level communication (e.g., NCCL tuning, InfiniBand, or RoCE v2) to minimize latency for 3D Gaussian Splatting (3DGS) and large-scale training.
  • Partner with researchers to fine-tune performance across multi-node GPU clusters for FSDP and DeepSpeed workloads.

Requirements

  • 5+ years in Cloud Infrastructure, DevOps, or MLOps supporting high-scale compute environments.
  • Deep expertise in K8s, Helm, and container orchestration.
  • Strong background in Apache Airflow, Argo Workflows, MLFlow, and Terraform.
  • Practical experience supporting frameworks like Ray and PyTorch Distributed.
  • Proficiency in Python, Bash scripting, and a solid understanding of IAM/RBAC.

Nice to Have

  • Distributed Training Expertise: Deep understanding of FSDP, and DeepSpeed.
  • AI Agent Orchestration: Experience building Agentic Workflows (LangGraph, AutoGen) for infrastructure automation or data curation.
  • Advanced Protocols: Familiarity with Model Context Protocol (MCP) to connect AI agents with infrastructure tools.

Work Arrangement

On-site — Santa Clara, CA

Additional Information

  • This role is onsite 5 days a week at the Santa Clara, CA office.
Required Skills
KubernetesHelmPythonMCP
About company
Gatik AI
Gatik is the leader in autonomous middle-mile logistics, revolutionizing the B2B supply chain with its autonomous transportation-as-a-service (ATaaS) solution. The company focuses on short-haul, B2B logistics for Fortune 500 retailers and launched the world’s first fully driverless commercial transportation service with Walmart.
All jobs at Gatik AI Visit website
Job Details
Category infrastructure
Posted 14 hours ago