Responsibilities
- Manage GPU and accelerator resources in Kubernetes environments with efficient scheduling, resource isolation, and topology-aware placement for inference tasks
- Build and maintain model serving platforms using vLLM, Triton Inference Server, TGI, or custom stacks with optimized batching, caching, and routing
- Implement intelligent routing and load balancing across diverse accelerator types including NVIDIA GPUs and AWS Inferentia/Trainentia to reduce latency and boost utilization
- Develop autoscaling solutions that adjust inference compute capacity dynamically based on demand across production and experimental workloads
- Design production-grade deployment workflows for machine learning models with canary releases, A/B testing, version control, and rollback capabilities across regions
- Define and manage infrastructure-as-code using Terraform and Helm for GPU-enabled EKS clusters, including node configuration, spot instance policies, and accelerator networking
- Establish observability practices with monitoring of GPU usage, latency metrics, token throughput, and service level objectives for model endpoints
- Build CI/CD pipelines for model artifacts with containerization including CUDA dependencies, model registry integration, and automated performance testing
- Architect AWS infrastructure for machine learning, including EKS with GPU nodes, accelerated EC2 instances, S3 for model storage, high-speed networking, and secure IAM policies
- Lead cost optimization initiatives through right-sizing of accelerator instances, strategic use of spot instances for inference, and efficiency reporting across compute fleets
Compensation
Competitive salary and equity package commensurate with experience
Work Arrangement
Hybrid remote with team coordination across time zones
Team
Collaborative engineering team focused on AI infrastructure and scalable machine learning systems
Requirements
- Extensive experience with Kubernetes and container orchestration in production environments
- Deep knowledge of GPU and accelerator architectures and their integration with container platforms
- Proven track record building scalable model serving systems for large language models
- Strong proficiency in Terraform, Helm, and infrastructure-as-code practices
- Hands-on experience with AWS services including EKS, EC2, S3, and EFA networking
- Familiarity with CI/CD pipelines and automated testing for machine learning models
- Experience with observability tools and performance profiling for GPU workloads
- Understanding of security best practices in cloud environments, including IAM and network policies
- Ability to optimize system efficiency and reduce cloud infrastructure costs
- Excellent problem-solving skills and ability to work across technical domains
Preferred Qualifications
- Experience with vLLM, Triton Inference Server, or similar model serving frameworks
- Background in distributed systems and low-latency service design
- Knowledge of CUDA, TensorRT, or other GPU compute stacks
- Contributions to open-source projects related to ML infrastructure
- Familiarity with multi-region deployments and global traffic management
Available for qualified candidates