Перейти к содержимому

GLM-5.3-Flash: 1,341/44 tok/s on 2× DGX Spark + Dual-Path patch

AmesianX

0:00 / 0:00

GLM-5.3-Flash: 1,341/44 tok/s on 2× DGX Spark + Dual-Path patch

27 просмотров · 2 дня назад
AmesianX
38 подписчиков
27 просмотров · 2 дня назад
Two DGX Sparks. A 320B MoE with a 1M-token context. 1,341 tok/s prefill and 44 tok/s decode — at the same time, from the same weights. DGX Spark 두 대. 320B MoE, 1M 컨텍스트. 프리필 1,341 tok/s 와 디코드 44 tok/s 를 같은 가중치에서 동시에. 보통은 둘 중 하나를 포기해야 합니다. ---------------------------------------------- | THE TRADE-OFF / 포기해야 했던 것 EN —Turning on FP8 for the dense layers buys ~40% decode throughput and costs ~12% of your prefill: the Marlin W8A16 kernel falls off at large M. The patch in this video stops paying that tax. Dense calls above M 1024 unpack Marlin FP8 back to BF16 (Triton, one pass) and run cuBLAS; decode calls at M 20 or below keep the Marlin kernel. One set of weights, two code paths, no trade-off. KO — dense 층에 FP8 을 켜면 디코드가 약 40% 빨라지는 대신 프리필을 약 12% 잃습니다. Marlin W8A16 커널이 큰 M 에서 효율이 떨어지기 때문입니다. 이 영상의 패치는 그 대가를 치르지 않습니다. M 이 1024 를 넘는 dense 호출은 Marlin FP8 을 BF16 으로 역변환해 (Triton 1패스) cuBLAS 로 태우고, M 이 20 이하인 디코드는 Marlin 을 그대로 씁니다. 가중치 하나, 경로 둘, 트레이드오프 없음. ---------------------------------------------- | THE NUMBERS / 실측치 Prefill 프리필 12k ctx warm 웜 1,341 tok/s 24k ctx warm 웜 1,311 tok/s first request cold 콜드 851 tok/s Decode 디코드 code 코드 43.2 tok/s Korean prose 한국어 산문 36.5 tok/s English prose 영문 산문 32.8-36.1 tok/s step time 스텝 ~82 ms Load to READY 로딩 225 s KV cache KV 캐시 847,533 tokens Max context served 최대 컨텍스트 524,288 tokens Why the dual path matters / 이중경로가 하는 일 prefill 12k decode prose all BF16 전부 BF16 1,395 25.6 dense FP8 only FP8만 1,225 36.2 ← +40% / -12% FP8 + dual-path 이중경로 1,341 36.5 ← both 둘 다 ---------------------------------------------- | THE SETUP / 구성 Hardware 하드웨어 2× NVIDIA DGX Spark (GB10), unified memory Interconnect 인터커넥트 200GbE ConnectX-7, RoCE v2, NCCL over IB verbs Model 모델 GLM-5.3-Flash — 320B MoE, 18B active, 1M context Quant 양자화 EXL3 TR3 4bpw Runtime 런타임 vLLM (EXL3 + B12X fork), tensor parallel across 2 nodes Spec decode 스펙 디코딩 native MTP, 4 tokens Precision, layer by layer / 층별 정밀도 MoE routed experts 라우팅 전문가 ...... W4A16 (EXL3 trellis, 4-bit weights) Dense + shared expert + lm_head ..... W8A16 (FP8 weights, Marlin) KV cache KV 캐시 .................... FP8 Activations 활성화 .................. BF16 everywhere / 전부 BF16 EN — That last line is deliberate. Quantizing activations on this model produced visible output corruption in our testing, so: weights only, activations untouched. No W4A8, no INT4. KO — 마지막 줄은 의도한 선택입니다. 이 모델에서 활성화를 양자화하니 출력이 눈에 띄게 망가졌습니다. 그래서 가중치만 깎고 활성화는 손대지 않았습니다. W4A8 도 INT4 도 없습니다. ---------------------------------------------- | MODELS & CREDIT / 모델과 출처 The exact checkpoint served here, with full lineage: 이 영상에서 실제로 서빙한 체크포인트와 그 계보입니다. Base 원본 zai-org/GLM-5.3-Flash huggingface.co/zai-org/GLM-5.3-Flash EXL3 TR3 4bpw quantization / 양자화 brandonmusic/GLM-5.3-Flash-tr3-4bpw huggingface.co/brandonmusic/GLM-5.3-Flash-tr3-4bpw Served here 실제 서빙 — abliterated (uncensored) variant of the above, produced by direct o_proj weight-space editing on all 45 decoder layers. 위 양자화본의 abliterated(무검열) 변형. 45개 디코더 층의 o_proj 를 가중치 공간에서 직접 편집한 것으로, 양자화된 전문가 텐서는 원본과 바이트 동일입니다. lovesenko/GLM-5.3-Flash-tr3-4bpw-Abliterated huggingface.co/lovesenko/GLM-5.3-Flash-tr3-4bpw-Abliterated Runtime image / 런타임 이미지 ghcr.io/entrpi/glm-5.3-flash-exl3-2x-spark:v2.3-tier1 github.com/Entrpi/glm-5.3-flash-exl3-2x-spark The dual-path patch is ours, layered on top of that image as an overlay. 이중경로 패치는 저희 것이며, 위 이미지 위에 오버레이로 얹습니다. ---------------------------------------------- | FOR CONTEXT / 비교 기준 EN — A published 4-node TP4 EXL3 reference on this same model reports 31.7 tok/s single-stream. This is two nodes at 36-44. huggingface.co/cfontes/GLM-5.3-Flash-EXL3-TP4-Spark KO — 같은 모델을 4노드 TP4 로 돌린 공개 기준치가 단일 스트림 31.7 tok/s 입니다. 이 구성은 2노드로 36~44 입니다. ---------------------------------------------- | CHAPTERS / 챕터 0:00 The trade-off nobody talks about / 아무도 말 안 하는 트레이드오프 0:00 Two DGX Sparks, 200G RoCE / 하드웨어 0:00 Why Marlin FP8 kills your prefill / FP8 이 프리필을 죽이는 이유 0:00 The dual-path patch / 이중경로 패치 0:00 Benchmarks / 실측 0:00 What's still on the table / 남은 여지 #DGXSpark #vLLM #LLM #GB10 #LocalLLM #Quantization #MoE #EXL3