Перейти к содержимому

LFM2.5-DSpark's 3.2x is a bytes-per-token story

Kemmu Draws Tech

0:00 / 0:00

LFM2.5-DSpark's 3.2x is a bytes-per-token story

52 просмотра · 9 дней назад
Kemmu Draws Tech
9 подписчиков
52 просмотра · 9 дней назад
Liquid AI publishes its DSpark speedup as a multiple rather than as a rate against a hardware limit. This episode rebuilds the memory-bandwidth roofline for single-stream decode, computes the bytes moved per token at bfloat16, and shows which candidate explanations for 3.18x can arithmetically produce it and which cannot — in particular that a ~300M drafter run nine times in front of a 1.2B target tops out at 1.47x. Chapters 0:00 A multiple with no units 0:28 The rates are printed 1:01 The bus and the truck 1:53 Bytes per token 2:37 The wrong two numbers 3:47 What speculation can change 5:00 The drafter is not free 5:42 Nine passes of a quarter-sized model 6:28 What the Markov head is for 7:36 The laptop column 8:10 The whole claim in one line Sources Up to 3.2x Faster Inference with LFM2.5-DSpark — https://huggingface.co/blog/LiquidAI/... Personal AI Supercomputer Powered by Blackwell | NVIDIA DGX Spark — https://www.nvidia.com/en-us/products... LiquidAI/LFM2-1.2B · Hugging Face — https://huggingface.co/LiquidAI/LFM2-... LiquidAI/LFM2-350M · Hugging Face — https://huggingface.co/LiquidAI/LFM2-... Fast Inference from Transformers via Speculative Decoding — https://arxiv.org/abs/2211.17192 Drawn with Inkstack.