Forget Benchmark Scores: Chinese AI Players Question the Value of Leaderboards

·
By Raisink Team

The world of artificial intelligence has long been obsessed with benchmark scores. Models are pitted against each other in standardized tests, and a few points higher on any given leaderboard can make or break a model’s reputation. However, at this year’s World AI Conference (WAIC) in Shanghai, Chinese AI players are pushing back against the assumption that these scores tell the whole story.

At WAIC 2026, companies like StepFun and MiniMax showcased their models’ capabilities beyond benchmark scores. Yang Minghui from StepFun emphasized that ‘the benchmark race is not the whole point.’ Instead of focusing on leaderboard rankings, his company prioritizes whether its model actually works well in real-world applications â phones, cars, and robots.

StepFun’s agent operating system STEPX Neo was one such example. The company argues that benchmark homogeneity is particularly problematic for companies like theirs that run multiple vertical-specific models. ‘Each of our models excels in different domains,’ Yang added, ‘no single benchmark can capture that diversity.’

This critique comes as the industry confronts a well-documented problem: benchmark contamination and saturation. As evaluation platform Singularity Moments noted, ‘benchmark contamination is a persistent issue’ and ‘differences within a few points [on MMLU] are often meaningless’ given how easily top models can memorize test data.

MiniMax’s Bai Chuanxu made the case from a different angle. The Shanghai-based firm’s flagship model M3 competes efficiently despite having fewer parameters compared to some rivals. ‘Our model supports images, video and audio natively,’ he explained, ‘but benchmark scores often fail to reflect that and can even create misleading impressions.’

Baidu, whose universal agent DuMate was named a WAIC 2026 official ‘treasure of the hall,’ also sees limits in conventional evaluation. Li Jingqiu, product lead for DuMate, put it differently: ‘The best way to judge a model today is to give it a complex, time-consuming task.’

Baidu’s founder Robin Li suggested that the industry should consider daily active agents as a score, just like daily active users in the age of web. This pragmatic turn â AI is not a monolith â unites these Chinese voices: there’s no one-size-fits-all score.

When models are deployed in real products serving millions of users, the metric that matters is whether the model actually delivers. As Yang Minghui from StepFun said, ‘The real test is just using it.’ This shift away from benchmark scores and toward practical evaluation marks a significant change in how Chinese AI players approach their work.

The industry’s reliance on standardized tests has led to concerns about data leakage and overfitting. Industry analyst site 80AJ observed that ‘static benchmarks are rapidly losing value due to these issues,’ forcing a shift toward human-preference arenas like Chatbot Arena.

Related news