Frontier labs continue pushing large language model performance on the Chatbot Arena benchmark through rapid releases and post-training refinements. As of late September 2026, Anthropic’s Claude Opus and Fable variants lead with Elo scores near 1505–1531, followed closely by Google’s Gemini 3.x series, OpenAI’s GPT-5.6 and GPT-6 models, Meta’s Muse Spark, and strong open-weight entries from Moonshot and Alibaba. Multiple new model drops in September, including Gemini 3.8 Flash and GPT-6 variants, demonstrate ongoing capability gains in reasoning, agentic tasks, and human preference. With roughly three months remaining, the race hinges on whether scaling, architectural tweaks, or specialized fine-tuning can lift any model past the implied threshold before December 31.
基於Polymarket數據的AI實驗性摘要。這不是交易建議,也不影響該市場的結算方式。 · 更新於$217,603 交易量
↑ 1530
37%
↑ 1540
15%
↑ 1550
17%
↑ 1600
5%
↑ 1650
3%
↑ 1700
4%
$217,603 交易量
↑ 1530
37%
↑ 1540
15%
↑ 1550
17%
↑ 1600
5%
↑ 1650
3%
↑ 1700
4%
Results from the 'Score' section on the 'Text Arena' Leaderboard tab (https://lmarena.ai/leaderboard/text), with the style control unchecked, will be used to resolve this market.
The resolution source is the Chatbot Arena LLM Leaderboard (https://lmarena.ai/). If this source is temporarily unavailable, the market remains open until it is accessible again; if permanently unavailable, this market will resolve to "No".
市場開放時間: Jul 23, 2026, 5:38 PM ET
Results from the 'Score' section on the 'Text Arena' Leaderboard tab (https://lmarena.ai/leaderboard/text), with the style control unchecked, will be used to resolve this market.
The resolution source is the Chatbot Arena LLM Leaderboard (https://lmarena.ai/). If this source is temporarily unavailable, the market remains open until it is accessible again; if permanently unavailable, this market will resolve to "No".
Frontier labs continue pushing large language model performance on the Chatbot Arena benchmark through rapid releases and post-training refinements. As of late September 2026, Anthropic’s Claude Opus and Fable variants lead with Elo scores near 1505–1531, followed closely by Google’s Gemini 3.x series, OpenAI’s GPT-5.6 and GPT-6 models, Meta’s Muse Spark, and strong open-weight entries from Moonshot and Alibaba. Multiple new model drops in September, including Gemini 3.8 Flash and GPT-6 variants, demonstrate ongoing capability gains in reasoning, agentic tasks, and human preference. With roughly three months remaining, the race hinges on whether scaling, architectural tweaks, or specialized fine-tuning can lift any model past the implied threshold before December 31.
基於Polymarket數據的AI實驗性摘要。這不是交易建議,也不影響該市場的結算方式。 · 更新於



警惕外部連結哦。
警惕外部連結哦。
Frequently Asked Questions