Hugging Face now has 48 official benchmarks. Here is what the map looks like
Hugging Face marks a growing set of datasets as official benchmarks . Each one gets a leaderboard on its dataset page, filled automatically from the .eval_results files that model repositories publish. There are now 48 of them, from GPQA Diamond and MMLU-Pro to SWE-bench, Terminal-Bench, AIME 2026…
Read the full story at DEV Community — Machine Learning ↗
Timeline · 2 reports
- 2026-10-04 15:16 · DEV Community — Machine Learning
Hugging Face official benchmarks: the complete list (48) and how their leaderboards work - 2026-10-04 15:15 · DEV Community — Machine Learning
Hugging Face now has 48 official benchmarks. Here is what the map looks like
More stories
- Aleph-Alpha/Kolibri-1 · Hugging Face - 78B parameters. 3.46B active. Up to 1M tokens of context - Apache 2.0 — r/LocalLLaMA
- Hinton says AI already has subjective experience. I'm not convinced, and the Hugging Face breach doesn't change that — r/ArtificialInteligence
- The ultimate guide to multi-harness RL — r/huggingface
- llama, server: add /v1/systemone API (models: laya, julia-1, lev, openjev, kev) by ngxson · Pull Request #29818 · ggml-org/llama.cpp — r/LocalLLaMA
- I fine-tuned SmolVLM-500M into a lightweight Windows OS Agent (<8GB VRAM) Looking for feedback & ideas! [Weights on HuggingFace] — r/huggingface
- [D] I open-sourced 30,000 paired QR-Code Illusions with multi-decoder verification & robustness scores on Hugging Face (Free for ControlNet / LoRA training) — r/StableDiffusion
- A quick Minimax H3 news round-up - 2nd October 2026 — r/comfyui
- Bytedance release 4-step for Minimax-h3; DMAD: Distribution Matching as Adversarial Distillation — r/StableDiffusion
Get the daily brief of stories like this at 6:30 every morning →