Santa Clara, Canada On-site Full-time USD 180,000 – 240,000 / year

Gatik AI is hiring a Senior AI Infrastructure Engineer

Responsibilities

  • Architect and deploy high-performance AI platforms that support large-scale autonomous driving models
  • Bridge research and production environments to ensure scalability, reliability, and efficiency in AI systems
  • Enable researchers to scale advanced models like Vision-Language Architectures and World Models using distributed frameworks such as PyTorch Distributed and Ray Train
  • Optimize multi-GPU cluster performance through efficient model and data parallelism strategies on H100 and A100 hardware
  • Tune low-level networking protocols including NCCL, InfiniBand, and RoCE v2 to reduce communication latency during training and 3DGS workloads
  • Improve GPU utilization and cost efficiency using Kubernetes-native scheduling tools like NVIDIA GPU Operator and KubeFlow
  • Enhance inference pipelines by deploying optimized models via TensorRT, ONNX Runtime, and Triton Inference Server for real-time and batch scenarios
  • Build self-monitoring systems using AI agents (e.g., LangGraph, CrewAI, AutoGen) to detect and respond to hardware failures and NCCL timeouts automatically
  • Develop automated DevOps workflows using AI agents that review infrastructure code and recommend Kubernetes resource tuning based on model needs
  • Support the creation of autonomous data curation systems where AI agents identify, label, and validate critical edge cases from raw sensor data
  • Automate the end-to-end machine learning lifecycle using tools like MLFlow, Argo Workflows, and Kubernetes
  • Integrate experiment tracking and feature stores to maintain a reliable record of all model versions and training runs
  • Implement safe deployment patterns such as A/B testing, shadow deployments, and automated rollback procedures
  • Enforce infrastructure consistency and reproducibility through Infrastructure-as-Code using Terraform and Helm
  • Collaborate on scaling ETL workflows with Apache Airflow, Kafka, and Spark for efficient data processing
  • Work with data engineering teams to build high-throughput data pipelines that move sensor data to cloud storage like S3, GCS, or Delta Lake
  • Define and monitor key performance indicators such as training speed, latency, throughput, and model drift
  • Ensure comprehensive system visibility using monitoring tools including Prometheus, Grafana, OpenTelemetry, and the ELK Stack
  • Establish deep observability across both infrastructure layers and ML metrics like convergence and throughput
  • Track and alert on AI-specific KPIs including inference latency, throughput, and feature drift

Work Arrangement

On-site — Santa Clara, CA

Other

This position requires full-time, in-person attendance from Monday to Friday at the Santa Clara, CA office.

About company
Gatik AI
Gatik is the leader in autonomous middle-mile logistics, revolutionizing the B2B supply chain with its autonomous transportation-as-a-service (ATaaS) solution. The company focuses on short-haul, B2B logistics for Fortune 500 retailers and launched the world’s first fully driverless commercial transportation service with Walmart.
All jobs at Gatik AI Visit website
Job Details
Category DevOps & SRE
Posted a day ago