Business

AI Benchmarks Are Breaking as Open Models Close the Gap

Why standard tests can no longer distinguish the strongest AI models

LONDON — “It’s getting hard to tell what the best model is,” said Thomas Wolf, co-founder and chief scientist of the open-source platform Hugging Face. Speaking at the Brainstorm AI conference in London, Wolf pointed to the marginal differences among recent flagship releases: “They all seem to be, actually, very close.”

Advertisement

The difficulty comes as top-tier models from Google, OpenAI, and Meta reach near-parity on standard academic tests. For years, the artificial intelligence ecosystem relied on standardized benchmarks to rank models, including GLUE (General Language Understanding Evaluation), introduced in 2018 to measure natural language understanding, and HellaSwag, a 2019 test designed to evaluate commonsense reasoning.

MMLU, or Massive Multitask Language Understanding, became the industry gold standard after its release in 2020 by researchers at UC Berkeley and other institutions. It comprises multiple-choice questions across 57 subjects, ranging from elementary mathematics to professional law. Contemporary models now routinely achieve near-perfect scores, however. “These benchmarks are mostly all saturated right now,” Wolf explained.

A February 2024 study from the European Commission’s Joint Research Centre, titled *“Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation,”* identified deep-seated flaws in current assessment practices. Its findings included misaligned commercial incentives, construct-validity failures, and data contamination, in which test questions accidentally leak into the massive datasets used to train models and artificially inflate their scores. “Just because these models are all working the same on this academic benchmark doesn’t really mean that they’re all exactly the same,” Wolf said.

The pressure to redesign evaluation frameworks has grown with the democratization of high-performing open-source AI. Silicon Valley’s closed-source giants historically dominated the landscape, but the balance of power shifted with DeepSeek, an AI research firm based in Hangzhou, China. Its R1 model was built at a fraction of the cost of its American counterparts and matched or exceeded proprietary models on reasoning tasks.

Wolf compared DeepSeek’s arrival with the debut of OpenAI’s flagship chatbot in late 2022. “Just like ChatGPT was the moment the whole world discovered AI, DeepSeek was the moment the whole world discovered there was kind of this open society,” he said, calling it a “ChatGPT moment” for open-source AI.

That movement is central to Hugging Face, which Wolf founded in 2016 with Clément Delangue and Julien Chaumond. The company began as an entertainment chatbot app for teenagers before pivoting to host open-source machine learning models. It is now widely considered the “GitHub of machine learning,” serving as a massive repository where developers share and deploy models, datasets, and applications.

Hugging Face was valued at $4.5 billion in its last major funding round in August 2023, which included backing from Google, Nvidia, and Amazon. Wolf emphasized that the company’s business model is fundamentally aligned with the open-source community and aims to maximize participation and model sharing.

Wolf predicted that the industry will move toward two distinct evaluation paradigms in 2025. One will focus on “agentic” capabilities: models will be assessed on whether they can act as autonomous agents by executing multi-step digital workflows, using external software tools, and writing code to solve complex problems.

The other paradigm will center on hyper-customized, use-case-specific evaluation. Hugging Face has introduced a program called “Your Bench” to lead this transition. Users upload a small set of internal documents, and the tool automatically generates a bespoke benchmark tailored to the operational tasks they require. Organizations can then compare how different models perform on their actual proprietary workflows rather than relying on abstract academic scores.

Related Articles

Leave a Reply

Your email address will not be published. Required fields are marked *