Tsinghua University introduces TokenRouter, boosting throughput up to 64X
Summary
A recent paper from Tsinghua University introduces TokenRouter, an innovative serving system that drastically improves the efficiency of token-level routing for large language models (LLMs), achieving up to 64.15 times the throughput of existing frameworks. Traditional serving frameworks, such as vLLM and SGLang, process one model per request, leading to synchronization delays when multiple models contribute to a single output. TokenRouter addresses this issue by allowing each model to operate on its own server, enabling them to collaborate more effectively while managing requests to optimize batch sizes and minimize idle time during inference.
Analysis
Tsinghua: Tsinghua University is a leading Chinese research institution with prominent programs in computer science and artificial intelligence. Researchers there recently published an arXiv paper introducing TokenRouter as a new serving system for token-level routing between small and large language models. TokenRouter: TokenRouter is a serving system built for per-token routing in LLM inference that assigns separate servers to different models and passes partial responses between them while preserving KV cache state. It was developed to overcome limitations in popular frameworks like vLLM and SGLang where a single model per request creates bottlenecks when responses involve multiple models. LLM Serving: Standard serving frameworks process one model per request, causing synchronization delays when multiple models contribute to a single response. Routing Efficiency: Token-level routing enables models to collaborate on responses while supporting larger batch sizes and reduced idle time during inference.
Categories
machine_learningaitech