TriForce is a hierarchical speculative decoding system designed for efficient serving of large language models (LLMs) in long contexts. It addresses the bottlenecks of KV cache and model weights, resulting in significant speedups.

5m read timeFrom marktechpost.com
Post cover image
1 Impression