Responsibilities
- Architect and manage distributed training systems for neural operator models such as Transolver and Point Cloud Transformer using NVIDIA DGX B200 hardware.
- Enhance training pipeline efficiency by improving throughput, fault tolerance, and cost-effectiveness through techniques like gradient accumulation and multi-node synchronization.
- Develop robust experiment tracking and observability tools to provide real-time insights into training progress, hyperparameter tuning, and model metrics.
- Address performance limitations in data loading for large-scale mesh-based datasets.
- Streamline data pipeline I/O operations from cloud storage via prefetching, caching, and optimized data formats.
- Handle integration of diverse data sources with varying formats, structures, and resolutions.
- Create scalable serving systems for pre-trained large physics models supporting zero-shot inference and uncertainty estimation using Monte Carlo Dropout.
- Develop automated pipelines for packaging models for external deployment.
- Ensure deployed models operate consistently in client environments and support fine-tuning workflows.
- Guarantee full reproducibility of model behavior from any saved checkpoint.
- Enhance research team productivity with tools enabling rapid iteration, stable CI/CD pipelines, and effective debugging capabilities.
- Partner with infrastructure teams across the organization to align on shared standards and reusable patterns.