Responsibilities
- Architect and deploy high-performance AI platforms that support large-scale autonomous driving models
- Bridge research and production environments to ensure scalability, reliability, and efficiency in AI systems
- Enable researchers to scale advanced models like Vision-Language Architectures and World Models using distributed frameworks such as PyTorch Distributed and Ray Train
- Optimize multi-GPU cluster performance through efficient model and data parallelism strategies on H100 and A100 hardware
- Tune low-level networking protocols including NCCL, InfiniBand, and RoCE v2 to reduce communication latency during training and 3DGS workloads
- Improve GPU utilization and cost efficiency using Kubernetes-native scheduling tools like NVIDIA GPU Operator and KubeFlow
- Enhance inference pipelines by deploying optimized models via TensorRT, ONNX Runtime, and Triton Inference Server for real-time and batch scenarios
- Build self-monitoring systems using AI agents (e.g., LangGraph, CrewAI, AutoGen) to detect and respond to hardware failures and NCCL timeouts automatically
- Develop automated DevOps workflows using AI agents that review infrastructure code and recommend Kubernetes resource tuning based on model needs
- Support the creation of autonomous data curation systems where AI agents identify, label, and validate critical edge cases from raw sensor data
- Automate the end-to-end machine learning lifecycle using tools like MLFlow, Argo Workflows, and Kubernetes
- Integrate experiment tracking and feature stores to maintain a reliable record of all model versions and training runs
- Implement safe deployment patterns such as A/B testing, shadow deployments, and automated rollback procedures
- Enforce infrastructure consistency and reproducibility through Infrastructure-as-Code using Terraform and Helm
- Collaborate on scaling ETL workflows with Apache Airflow, Kafka, and Spark for efficient data processing
- Work with data engineering teams to build high-throughput data pipelines that move sensor data to cloud storage like S3, GCS, or Delta Lake
- Define and monitor key performance indicators such as training speed, latency, throughput, and model drift
- Ensure comprehensive system visibility using monitoring tools including Prometheus, Grafana, OpenTelemetry, and the ELK Stack
- Establish deep observability across both infrastructure layers and ML metrics like convergence and throughput
- Track and alert on AI-specific KPIs including inference latency, throughput, and feature drift
Work Arrangement
On-site — Santa Clara, CA
Other
This position requires full-time, in-person attendance from Monday to Friday at the Santa Clara, CA office.