Responsibilities
- Manage and expand the GPU fleet to maximize useful inference throughput per GPU-dollar while meeting latency and reliability SLOs, and continuously reduce cost per token.
- Enable serverless inference for the open-weight model catalog and dedicated, tuned model deployments for enterprises.
- Build serverless serving pools using vLLM or SGLang engines with continuous batching, prefix caching, low-precision serving (FP8, FP4, INT4), and MoE expert parallelism for multi-tenant scale.
- Develop the fleet layer with llm-d or NVIDIA Dynamo on the Kubernetes Gateway API, including KV-cache-aware routing, prefill/decode disaggregation, KV tiering (Mooncake), and multi-LoRA serving for shared base pools.
- Construct the underlying platform with Kubernetes on bare metal, GPU Operator, topology-aware scheduling, LeaderWorkerSet, Kueue, Kata Containers for isolated tenants, and bare-metal lifecycle via OpenStack Ironic.
- Implement weight logistics and elasticity using P2P model distribution (Dragonfly, safetensors streaming), warm pools, and autoscaling driven by inference metrics such as queue depth, KV occupancy, and TTFT.
- Build the dedicated tier with per-tenant pools, GPU-hour metering, latency SLOs, private networking, and support for the model-tuning loop including adapter versioning, canary rollout, and rollback on outcome regression.
- Create observability and economics systems by integrating DCGM and engine metrics into Prometheus or OpenTelemetry, and perform capacity planning using roofline math for bandwidth-bound decode, batching curves, and utilization versus cost per token.
Benefits
- Access to the newest hardware, from the metal up, with a globally expanding B300 fleet built and operated end to end.
- Own the platform and build the team, with a greenfield inference platform where you set the pattern others follow.
- Open-source first approach, with upstream contribution as part of the job and supported conference travel.
- Based in mainland China with no relocation required, working remotely with a global, async-friendly team.
- If you choose to move, hiring entities exist in the Netherlands, the United States, Ireland, and Saudi Arabia, with visa sponsorship available.
Work Arrangement
Remote (Worldwide) — mainland China
What you'll build
- Serverless serving pools using vLLM or SGLang engines with continuous batching, prefix caching, low-precision serving (FP8, FP4, INT4), and MoE expert parallelism for multi-tenant scale.
- The fleet layer with llm-d or NVIDIA Dynamo on the Kubernetes Gateway API, including KV-cache-aware routing, prefill/decode disaggregation, KV tiering (Mooncake), and multi-LoRA serving for shared base pools.
- The platform under it with Kubernetes on bare metal, GPU Operator, topology-aware scheduling, LeaderWorkerSet, Kueue, Kata Containers for isolated tenants, and bare-metal lifecycle via OpenStack Ironic.
- Weight logistics and elasticity using P2P model distribution (Dragonfly, safetensors streaming), warm pools, and autoscaling driven by inference metrics such as queue depth, KV occupancy, and TTFT.
- The dedicated tier with per-tenant pools, GPU-hour metering, latency SLOs, private networking, and support for the model-tuning loop including adapter versioning, canary rollout, and rollback on outcome regression.
- Observability and economics by integrating DCGM and engine metrics into Prometheus or OpenTelemetry, and performing capacity planning using roofline math for bandwidth-bound decode, batching curves, and utilization versus cost per token.
The stack you'll work in
- Some of this is committed direction: Kubernetes on bare metal, vLLM or SGLang, and an OpenAI-compatible endpoint. Much of the rest is candidates you'll evaluate. You'll select, benchmark, and integrate components that earn their place in production. We value depth in the core serving stack and sound architectural judgment, not prior experience with every project listed.
- Engines: vLLM, SGLang
- Serving techniques: prefill/decode disaggregation, wide expert parallelism, speculative decoding (MTP, EAGLE-3), low precision (FP8, FP4, INT4)
- MoE and attention libraries: DeepEP, DeepGEMM, EPLB, FlashMLA
- Routing and serving: llm-d, NVIDIA Dynamo, Kubernetes Gateway API, Envoy, Ray Serve, KubeRay
- KV cache and weights: Mooncake, HiCache, 3FS, Dragonfly
- Scheduling and GPU sharing: Kubernetes, Volcano, Kueue, LeaderWorkerSet, HAMi
- Isolation and bare metal: Kata Containers, OpenStack Ironic, NVIDIA GPU Operator
- Observability: Prometheus, Grafana, DCGM, OpenTelemetry
When you apply
Include a brief description of an inference system you personally improved: the bottleneck, your intervention, and the measured result. An anonymized example is welcome.
Other
- Language requirements: English and Chinese
- Remote work: based in mainland China, no relocation required
- Visa sponsorship: available for relocation to Netherlands, United States, Ireland, Saudi Arabia
Available for relocation to Netherlands, United States, Ireland, Saudi Arabia