Sequoia is a scalable, robust, and hardware-aware speculative decoding framework that enables serving LLMs on consumer GPUs with low latency. It can serve a Llama2-70B on a single RTX-4090 8 times faster than other offloading serving systems. Sequoia is scalable and robust, allowing for faster growth in accepted tokens and generating temperatures more effectively compared to alternative methods.

1m read timeFrom infini-ai-lab.github.io
Post cover image
44 Impressions