LFM2.5-DSpark's 3.2x is a bytes-per-token story
Kemmu Draws Tech
0:00 / 0:00
LFM2.5-DSpark's 3.2x is a bytes-per-token story
52 просмотра · 9 дней назад
Kemmu Draws Tech
9 подписчиков
52 просмотра · 9 дней назад
Liquid AI publishes its DSpark speedup as a multiple rather than as a rate against a hardware limit. This episode rebuilds the memory-bandwidth roofline for single-stream decode, computes the bytes moved per token at bfloat16, and shows which candidate explanations for 3.18x can arithmetically produce it and which cannot — in particular that a ~300M drafter run nine times in front of a 1.2B target tops out at 1.47x.
Chapters
0:00 A multiple with no units
0:28 The rates are printed
1:01 The bus and the truck
1:53 Bytes per token
2:37 The wrong two numbers
3:47 What speculation can change
5:00 The drafter is not free
5:42 Nine passes of a quarter-sized model
6:28 What the Markov head is for
7:36 The laptop column
8:10 The whole claim in one line
Sources
Up to 3.2x Faster Inference with LFM2.5-DSpark — https://huggingface.co/blog/LiquidAI/...
Personal AI Supercomputer Powered by Blackwell | NVIDIA DGX Spark — https://www.nvidia.com/en-us/products...
LiquidAI/LFM2-1.2B · Hugging Face — https://huggingface.co/LiquidAI/LFM2-...
LiquidAI/LFM2-350M · Hugging Face — https://huggingface.co/LiquidAI/LFM2-...
Fast Inference from Transformers via Speculative Decoding — https://arxiv.org/abs/2211.17192
Drawn with Inkstack.