Responsibilities
- Implement and productionize optimization techniques including quantization, speculative decoding, KV cache optimization, continuous batching, and LoRA serving
- Deep dive into inference frameworks (vLLM, SGLang, TensorRT-LLM) and underlying libraries to debug and improve performance
- Profile and optimize CUDA kernels and GPU utilization across our serving infrastructure
- Add support for new model architectures, ensuring they meet our performance standards before going to production
- Experiment with novel inference techniques and bring successful approaches into production
- Build tooling and benchmarks to measure and track inference performance across our fleet
- Collaborate with applied ML engineers to ensure trained models can be served efficiently
Requirements
- 2+ years of experience in ML systems, inference optimization, or GPU programming
- Strong proficiency in Python and familiarity with C++
- Hands-on experience with LLM inference frameworks (vLLM, SGLang, TensorRT-LLM, or similar)
- Deep understanding of GPU architecture and experience profiling GPU workloads
- Familiarity with LLM optimization techniques (quantization, speculative decoding, continuous batching, KV cache management)
- Experience with PyTorch and understanding of how models execute on hardware
- Track record of measurably improving system performance
Nice to Have
- Experience with CUDA programming
- Familiarity with serving non-LLM models (TTS, vision, embeddings)
- Experience with distributed inference and multi-GPU serving
- Contributions to open-source inference frameworks
- Experience with Docker and Kubernetes
Benefits
- Competitive compensation
- Equity in a high-growth startup
- Comprehensive benefits
Compensation
Competitive compensation, equity in a high-growth startup, and comprehensive benefits. The base salary range for this role is $220,000 - $320,000, plus equity and benefits, depending on experience.
Work Arrangement
Hybrid — San Francisco
Additional Information
- Most team members are in the office 4 days a week in SF
- Hybrid work available for Bay Area candidates
- Equal opportunity employer; no discrimination based on race, color, religion, gender, sexual orientation, national origin, genetics, disability, age, or veteran status