San Francisco, United States of America Hybrid Full-time USD 220,000 – 320,000 / year

Inference is hiring a Senior Software Engineer - Model Performance

Responsibilities

  • Implement and productionize optimization techniques including quantization, speculative decoding, KV cache optimization, continuous batching, and LoRA serving
  • Deep dive into inference frameworks (vLLM, SGLang, TensorRT-LLM) and underlying libraries to debug and improve performance
  • Profile and optimize CUDA kernels and GPU utilization across our serving infrastructure
  • Add support for new model architectures, ensuring they meet our performance standards before going to production
  • Experiment with novel inference techniques and bring successful approaches into production
  • Build tooling and benchmarks to measure and track inference performance across our fleet
  • Collaborate with applied ML engineers to ensure trained models can be served efficiently

Requirements

  • 2+ years of experience in ML systems, inference optimization, or GPU programming
  • Strong proficiency in Python and familiarity with C++
  • Hands-on experience with LLM inference frameworks (vLLM, SGLang, TensorRT-LLM, or similar)
  • Deep understanding of GPU architecture and experience profiling GPU workloads
  • Familiarity with LLM optimization techniques (quantization, speculative decoding, continuous batching, KV cache management)
  • Experience with PyTorch and understanding of how models execute on hardware
  • Track record of measurably improving system performance

Nice to Have

  • Experience with CUDA programming
  • Familiarity with serving non-LLM models (TTS, vision, embeddings)
  • Experience with distributed inference and multi-GPU serving
  • Contributions to open-source inference frameworks
  • Experience with Docker and Kubernetes

Benefits

  • Competitive compensation
  • Equity in a high-growth startup
  • Comprehensive benefits

Compensation

Competitive compensation, equity in a high-growth startup, and comprehensive benefits. The base salary range for this role is $220,000 - $320,000, plus equity and benefits, depending on experience.

Work Arrangement

Hybrid — San Francisco

Additional Information

  • Most team members are in the office 4 days a week in SF
  • Hybrid work available for Bay Area candidates
  • Equal opportunity employer; no discrimination based on race, color, religion, gender, sexual orientation, national origin, genetics, disability, age, or veteran status
Required Skills
PythonPytorchDockerKubernetes
About company
Inference
We are building a real-time marketplace for AI inference that matches spare GPU capacity inside data centers with demand from developers building AI-powered applications. We currently operate the world's largest distributed GPU cluster, with over 5,000 GPUs, hundreds of individual operators, and millions of gigabytes of VRAM connected to the network at any given moment. We are a small, well-funded team working on difficult, high-impact problems at the intersection of AI and distributed systems. We primarily work in-person from our office in downtown San Francisco. Our investors include A16z CSX and Multicoin. We are high-agency, adaptable, and collaborative. We value creativity alongside technical prowess and humility. We work hard, and deeply enjoy the work that we do.
All jobs at Inference Visit website
Job Details
Department Engineering
Category other
Posted 6 months ago