Responsibilities
- Create and maintain backend services in Go or Python that support a managed infrastructure platform.
- Build and manage control plane components for deploying and overseeing Kubernetes and Slurm clusters on GPU-heavy environments.
- Develop distributed systems to automate cluster provisioning, workload scheduling, and infrastructure lifecycle tasks.
- Implement platform features using Kubernetes APIs, controllers, operators, and cloud-native tools.
- Enhance platform reliability, scalability, security, and monitoring through automation and operational improvements.
- Troubleshoot and resolve production issues involving Kubernetes, distributed systems, networking, and cloud infrastructure.
- Work closely with infrastructure, AI, and platform teams to influence cloud platform development.
- Participate in technical design, system architecture, mentoring, best practices, and on-call support.
Work Arrangement
Hybrid — San Francisco, NYC, Seattle
Other
- This role requires presence in a U.S. office hub (San Francisco, NYC, or Seattle, with preference for SF or NYC), including at least two in-office days weekly and periodic offsites.
- Visa sponsorship is not offered for this position.
Not available at this time