Section 01
nano-vllm-qos: A QoS-Aware Scheduling System for LLM Inference (Introduction)
nano-vllm-qos: A QoS-Aware Scheduling System for LLM Inference
Key Points: This project is an SLO-aware scheduling system based on nano-vLLM, integrating Radix prefix caching and Mooncake remote KV caching technologies to optimize latency and throughput for LLM inference.
Original Author & Source
- Original Author/Maintainer: Xuhang0607
- Source Platform: GitHub
- Original Link: https://github.com/Xuhang0607/nano-vllm-qos
- Release/Update Date: 2026-08-11
Keywords: LLM Inference, QoS, SLO Scheduling, Prefix Caching, KV Caching, vLLM, Mooncake, Inference Optimization