MLPerf Inference v6.1 Press Briefing Q3 2026
MLCommons
0:00 / 0:00
MLPerf Inference v6.1 Press Briefing Q3 2026
83 просмотра · 9 дней назад
MLCommons
428 подписчиков
83 просмотра · 9 дней назад
On September 10, 2026, MLCommons held the press briefing for MLPerf Inference v6.1. David Kanter (MLCommons founder and Head of MLPerf), Miro Hodak (Inference WG co-chair), Ramesh Chukka (End-to-End RAG task force chair), and Palanivel Guruva Reddiar (Agentic Edge Inference task force chair) presented the latest results and introduced two new benchmarks.
Highlights:
486 datacenter and edge results from 30 organizations
gpt-oss-120b is the most popular workload in MLPerf history
First End-to-End RAG benchmark - real-world retrieval pipelines
First Agentic Edge Inference benchmark - on-device agentic AI
2.7X-5.7X per-accelerator gains in one year
Cisco's first cross-vendor submission (NVIDIA + AMD)
Crusoe's 512-accelerator run, the largest ever
NVIDIA Vera Rubin NVL72 preview debut
Results: https://mlcommons.org/visualizer
Datacenter: https://mlcommons.org/benchmarks/infe...
Edge: https://mlcommons.org/benchmarks/infe...
#MLPerf #MLCommons #AI #MachineLearning #Benchmarks #Inference #LLM #RAG #AgenticAI