Llm-D
@llm-d.ai
Kubernetes-native, high-performance distributed LLM inference
Llm-D's Company Logos
Llm-D's Brand Colors
Hex Code
Color name
RGB
HSL
CMYK
#7F317F
Plum
127, 49, 127
300, 44, 35
0, 61, 0, 50
#121619
Woodsmoke
18, 22, 25
206, 16, 8
28, 12, 0, 90
#FFDD80
Salomie
255, 221, 128
44, 100, 75
0, 13, 50, 0
About Llm-D
It helps teams run inference engines such as vLLM and SGLang across a cluster, turning single-node workloads into production-grade services on existing hardware. Designed for portability, llm-d supports GPUs, Google TPUs, Intel XPUs, CPUs, and emerging NPUs, offering consistent deployment patterns across heterogeneous accelerators.
The project provides “Well-Lit Paths”: tested, adaptable deployment recipes for common production needs. These include optimized baselines, latency-aware request routing, prefix-cache management, prefill/decode disaggregation, expert parallelism for mixture-of-experts models, flow control, fairness, and inference-pool autoscaling. The recipes help teams get started and tailor deployments to different models, hardware, and workloads.
llm-d also shares benchmarks, architecture and setup documentation, release updates, and community resources. Its stated performance results include higher output throughput and faster time to first token, with outcomes varying by configuration. A CNCF Sandbox project, llm-d welcomes users and contributors through its GitHub and Slack communities and is released under the Apache 2.0 License.
Brand industry
Computers Electronics and Technology
Company type
Suggest company type
Year founded
Suggest founded year
Company size
Suggest company size
