GLM-5.3-Flash: 1,341/44 tok/s on 2× DGX Spark + Dual-Path patch
AmesianX
0:00 / 0:00
GLM-5.3-Flash: 1,341/44 tok/s on 2× DGX Spark + Dual-Path patch
27 просмотров · 2 дня назад
AmesianX
38 подписчиков
27 просмотров · 2 дня назад
Two DGX Sparks. A 320B MoE with a 1M-token context. 1,341 tok/s prefill and
44 tok/s decode — at the same time, from the same weights.
DGX Spark 두 대. 320B MoE, 1M 컨텍스트. 프리필 1,341 tok/s 와 디코드 44 tok/s 를
같은 가중치에서 동시에. 보통은 둘 중 하나를 포기해야 합니다.
----------------------------------------------
| THE TRADE-OFF / 포기해야 했던 것
EN —Turning on FP8 for the dense layers buys ~40% decode throughput and costs
~12% of your prefill: the Marlin W8A16 kernel falls off at large M. The patch in
this video stops paying that tax. Dense calls above M 1024 unpack Marlin FP8
back to BF16 (Triton, one pass) and run cuBLAS; decode calls at M 20 or below
keep the Marlin kernel. One set of weights, two code paths, no trade-off.
KO — dense 층에 FP8 을 켜면 디코드가 약 40% 빨라지는 대신 프리필을 약 12% 잃습니다.
Marlin W8A16 커널이 큰 M 에서 효율이 떨어지기 때문입니다. 이 영상의 패치는 그 대가를
치르지 않습니다. M 이 1024 를 넘는 dense 호출은 Marlin FP8 을 BF16 으로 역변환해
(Triton 1패스) cuBLAS 로 태우고, M 이 20 이하인 디코드는 Marlin 을 그대로 씁니다.
가중치 하나, 경로 둘, 트레이드오프 없음.
----------------------------------------------
| THE NUMBERS / 실측치
Prefill 프리필 12k ctx warm 웜 1,341 tok/s
24k ctx warm 웜 1,311 tok/s
first request cold 콜드 851 tok/s
Decode 디코드 code 코드 43.2 tok/s
Korean prose 한국어 산문 36.5 tok/s
English prose 영문 산문 32.8-36.1 tok/s
step time 스텝 ~82 ms
Load to READY 로딩 225 s
KV cache KV 캐시 847,533 tokens
Max context served 최대 컨텍스트 524,288 tokens
Why the dual path matters / 이중경로가 하는 일
prefill 12k decode prose
all BF16 전부 BF16 1,395 25.6
dense FP8 only FP8만 1,225 36.2 ← +40% / -12%
FP8 + dual-path 이중경로 1,341 36.5 ← both 둘 다
----------------------------------------------
| THE SETUP / 구성
Hardware 하드웨어 2× NVIDIA DGX Spark (GB10), unified memory
Interconnect 인터커넥트 200GbE ConnectX-7, RoCE v2, NCCL over IB verbs
Model 모델 GLM-5.3-Flash — 320B MoE, 18B active, 1M context
Quant 양자화 EXL3 TR3 4bpw
Runtime 런타임 vLLM (EXL3 + B12X fork), tensor parallel across 2 nodes
Spec decode 스펙 디코딩 native MTP, 4 tokens
Precision, layer by layer / 층별 정밀도
MoE routed experts 라우팅 전문가 ...... W4A16 (EXL3 trellis, 4-bit weights)
Dense + shared expert + lm_head ..... W8A16 (FP8 weights, Marlin)
KV cache KV 캐시 .................... FP8
Activations 활성화 .................. BF16 everywhere / 전부 BF16
EN — That last line is deliberate. Quantizing activations on this model produced
visible output corruption in our testing, so: weights only, activations untouched.
No W4A8, no INT4.
KO — 마지막 줄은 의도한 선택입니다. 이 모델에서 활성화를 양자화하니 출력이 눈에 띄게
망가졌습니다. 그래서 가중치만 깎고 활성화는 손대지 않았습니다. W4A8 도 INT4 도 없습니다.
----------------------------------------------
| MODELS & CREDIT / 모델과 출처
The exact checkpoint served here, with full lineage:
이 영상에서 실제로 서빙한 체크포인트와 그 계보입니다.
Base 원본
zai-org/GLM-5.3-Flash
huggingface.co/zai-org/GLM-5.3-Flash
EXL3 TR3 4bpw quantization / 양자화
brandonmusic/GLM-5.3-Flash-tr3-4bpw
huggingface.co/brandonmusic/GLM-5.3-Flash-tr3-4bpw
Served here 실제 서빙 — abliterated (uncensored) variant of the above,
produced by direct o_proj weight-space editing on all 45 decoder layers.
위 양자화본의 abliterated(무검열) 변형. 45개 디코더 층의 o_proj 를
가중치 공간에서 직접 편집한 것으로, 양자화된 전문가 텐서는 원본과 바이트 동일입니다.
lovesenko/GLM-5.3-Flash-tr3-4bpw-Abliterated
huggingface.co/lovesenko/GLM-5.3-Flash-tr3-4bpw-Abliterated
Runtime image / 런타임 이미지
ghcr.io/entrpi/glm-5.3-flash-exl3-2x-spark:v2.3-tier1
github.com/Entrpi/glm-5.3-flash-exl3-2x-spark
The dual-path patch is ours, layered on top of that image as an overlay.
이중경로 패치는 저희 것이며, 위 이미지 위에 오버레이로 얹습니다.
----------------------------------------------
| FOR CONTEXT / 비교 기준
EN — A published 4-node TP4 EXL3 reference on this same model reports 31.7 tok/s
single-stream. This is two nodes at 36-44.
huggingface.co/cfontes/GLM-5.3-Flash-EXL3-TP4-Spark
KO — 같은 모델을 4노드 TP4 로 돌린 공개 기준치가 단일 스트림 31.7 tok/s 입니다.
이 구성은 2노드로 36~44 입니다.
----------------------------------------------
| CHAPTERS / 챕터
0:00 The trade-off nobody talks about / 아무도 말 안 하는 트레이드오프
0:00 Two DGX Sparks, 200G RoCE / 하드웨어
0:00 Why Marlin FP8 kills your prefill / FP8 이 프리필을 죽이는 이유
0:00 The dual-path patch / 이중경로 패치
0:00 Benchmarks / 실측
0:00 What's still on the table / 남은 여지
#DGXSpark #vLLM #LLM #GB10 #LocalLLM #Quantization #MoE #EXL3