China - Remote Remote (Global) Full-time

Telnyx is hiring an Inference Infrastructure Architect (Remote)

Responsibilities

  • Manage and expand the GPU fleet to maximize useful inference throughput per GPU-dollar while meeting latency and reliability SLOs, and continuously reduce cost per token.
  • Enable serverless inference for the open-weight model catalog and dedicated, tuned model deployments for enterprises.
  • Build serverless serving pools using vLLM or SGLang engines with continuous batching, prefix caching, low-precision serving (FP8, FP4, INT4), and MoE expert parallelism for multi-tenant scale.
  • Develop the fleet layer with llm-d or NVIDIA Dynamo on the Kubernetes Gateway API, including KV-cache-aware routing, prefill/decode disaggregation, KV tiering (Mooncake), and multi-LoRA serving for shared base pools.
  • Construct the underlying platform with Kubernetes on bare metal, GPU Operator, topology-aware scheduling, LeaderWorkerSet, Kueue, Kata Containers for isolated tenants, and bare-metal lifecycle via OpenStack Ironic.
  • Implement weight logistics and elasticity using P2P model distribution (Dragonfly, safetensors streaming), warm pools, and autoscaling driven by inference metrics such as queue depth, KV occupancy, and TTFT.
  • Build the dedicated tier with per-tenant pools, GPU-hour metering, latency SLOs, private networking, and support for the model-tuning loop including adapter versioning, canary rollout, and rollback on outcome regression.
  • Create observability and economics systems by integrating DCGM and engine metrics into Prometheus or OpenTelemetry, and perform capacity planning using roofline math for bandwidth-bound decode, batching curves, and utilization versus cost per token.

Benefits

  • Access to the newest hardware, from the metal up, with a globally expanding B300 fleet built and operated end to end.
  • Own the platform and build the team, with a greenfield inference platform where you set the pattern others follow.
  • Open-source first approach, with upstream contribution as part of the job and supported conference travel.
  • Based in mainland China with no relocation required, working remotely with a global, async-friendly team.
  • If you choose to move, hiring entities exist in the Netherlands, the United States, Ireland, and Saudi Arabia, with visa sponsorship available.

Work Arrangement

Remote (Worldwide) — mainland China

What you'll build

  • Serverless serving pools using vLLM or SGLang engines with continuous batching, prefix caching, low-precision serving (FP8, FP4, INT4), and MoE expert parallelism for multi-tenant scale.
  • The fleet layer with llm-d or NVIDIA Dynamo on the Kubernetes Gateway API, including KV-cache-aware routing, prefill/decode disaggregation, KV tiering (Mooncake), and multi-LoRA serving for shared base pools.
  • The platform under it with Kubernetes on bare metal, GPU Operator, topology-aware scheduling, LeaderWorkerSet, Kueue, Kata Containers for isolated tenants, and bare-metal lifecycle via OpenStack Ironic.
  • Weight logistics and elasticity using P2P model distribution (Dragonfly, safetensors streaming), warm pools, and autoscaling driven by inference metrics such as queue depth, KV occupancy, and TTFT.
  • The dedicated tier with per-tenant pools, GPU-hour metering, latency SLOs, private networking, and support for the model-tuning loop including adapter versioning, canary rollout, and rollback on outcome regression.
  • Observability and economics by integrating DCGM and engine metrics into Prometheus or OpenTelemetry, and performing capacity planning using roofline math for bandwidth-bound decode, batching curves, and utilization versus cost per token.

The stack you'll work in

  • Some of this is committed direction: Kubernetes on bare metal, vLLM or SGLang, and an OpenAI-compatible endpoint. Much of the rest is candidates you'll evaluate. You'll select, benchmark, and integrate components that earn their place in production. We value depth in the core serving stack and sound architectural judgment, not prior experience with every project listed.
  • Engines: vLLM, SGLang
  • Serving techniques: prefill/decode disaggregation, wide expert parallelism, speculative decoding (MTP, EAGLE-3), low precision (FP8, FP4, INT4)
  • MoE and attention libraries: DeepEP, DeepGEMM, EPLB, FlashMLA
  • Routing and serving: llm-d, NVIDIA Dynamo, Kubernetes Gateway API, Envoy, Ray Serve, KubeRay
  • KV cache and weights: Mooncake, HiCache, 3FS, Dragonfly
  • Scheduling and GPU sharing: Kubernetes, Volcano, Kueue, LeaderWorkerSet, HAMi
  • Isolation and bare metal: Kata Containers, OpenStack Ironic, NVIDIA GPU Operator
  • Observability: Prometheus, Grafana, DCGM, OpenTelemetry

When you apply

Include a brief description of an inference system you personally improved: the bottleneck, your intervention, and the measured result. An anonymized example is welcome.

Other

  • Language requirements: English and Chinese
  • Remote work: based in mainland China, no relocation required
  • Visa sponsorship: available for relocation to Netherlands, United States, Ireland, Saudi Arabia

Available for relocation to Netherlands, United States, Ireland, Saudi Arabia

Job Details
Location China - Remote
Work mode Remote (Global)
Employment Full-time
Category other
Posted 11 days ago
or drop your CV first
About company
Telnyx logo

Telnyx delivers global, low-latency voice and messaging infrastructure powered by a private global network. The company provides Voice AI agents, programmable communications APIs, and eSIM solutions for businesses requiring high-quality, secure, and scalable real-time interactions.

With full ownership of its telecom stack—from carrier network to AI inference—Telnyx enables enterprises to deploy autonomous, real-time AI agents with unmatched call quality, compliance, and speed. Its platform supports use cases across healthcare, finance, travel, logistics, and more.

Telnyx is trusted by over 14,000 industry-leading companies, including OpenAI, IBM, Cisco, Microsoft, and Zillow, and complies with global standards such as SOC 2, HIPAA, GDPR, and PCI.

All jobs at Telnyx Visit website