Responsibilities
- Design, build, and maintain scalable control plane services, operators, and custom Kubernetes controllers; develop automation in Go/Python for end-to-end cluster lifecycle management — provisioning, upgrades, patching, and deletion
- Build GPU-aware orchestration systems, working within the platform architecture to support GPU scheduling and resource allocation
- Partner with the Network team on networking solutions for AI workloads: CNI integration (Cilium, Multus), high-performance fabrics (InfiniBand, RoCE), RDMA, and GPUDirect
- Write resilient systems that handle failure gracefully — timeouts, retries, backoff, and degraded-mode operation — across large-scale distributed environments
- Develop platform services for inference: model serving infrastructure, autoscaling based on inference load, and multi-model deployment patterns
- Build internal tools and CLIs that let ML/AI teams deploy and monitor their own inference services
- Support and debug production issues through on-call rotation
Benefits
- Generous cash & equity compensation
- Health, dental, and vision coverage for you and your dependents
- Wellness and commuter stipends for select roles
- 401k Plan with 2% company match (USA employees)
- Flexible paid time off plan that we all actually use
Compensation
Generous cash & equity compensation
Work Arrangement
Hybrid — San Francisco, San Jose, Bellevue
Other
- This position requires presence in our San Francisco, San Jose, or Bellevue office location 4 days per week; Lambda’s designated work from home day is currently Tuesday.
- Equal Opportunity Employer