Anthropic's Claude family, including the recently released Claude Fable 5 and Fable 5.1 variants, currently leads LMSYS-style Overall Arena leaderboards with Elo scores near 1506 as of mid-September 2026, reflecting strong human preference in blind evaluations across writing, reasoning, and coding tasks. OpenAI's GPT-6 Astra and Meta's Muse Spark 1.2/1.3 series have narrowed gaps through iterative improvements in large language model capabilities, while open-weight contenders like Qwen3.8 Max and Kimi K3 add competitive pressure. Trader sentiment hinges on whether frontier labs can deliver further gains via scaling, post-training, or new architectures before year-end, with historical patterns showing 10-30 point monthly advances amid tight races. Key catalysts include anticipated Q4 releases, developer conferences, and any major capability demonstrations that could push the top score higher.
Resumo experimental gerado por IA com dados do Polymarket. Isto não é aconselhamento de trading e não tem qualquer papel na resolução deste mercado. · AtualizadoAlgum modelo de IA alcançará ___ Pontuação Geral da Arena até 31 de dezembro?
$185,780 Vol.
↑ 1520
97%
↑ 1530
33%
↑ 1540
23%
↑ 1550
21%
↑ 1600
5%
↑ 1650
5%
↑ 1700
3%
$185,780 Vol.
↑ 1520
97%
↑ 1530
33%
↑ 1540
23%
↑ 1550
21%
↑ 1600
5%
↑ 1650
5%
↑ 1700
3%
Results from the 'Score' section on the 'Text Arena' Leaderboard tab (https://lmarena.ai/leaderboard/text), with the style control unchecked, will be used to resolve this market.
The resolution source is the Chatbot Arena LLM Leaderboard (https://lmarena.ai/). If this source is temporarily unavailable, the market remains open until it is accessible again; if permanently unavailable, this market will resolve to "No".
Mercado Aberto: Jul 23, 2026, 5:38 PM ET
Resolver
0x65070BE91...Results from the 'Score' section on the 'Text Arena' Leaderboard tab (https://lmarena.ai/leaderboard/text), with the style control unchecked, will be used to resolve this market.
The resolution source is the Chatbot Arena LLM Leaderboard (https://lmarena.ai/). If this source is temporarily unavailable, the market remains open until it is accessible again; if permanently unavailable, this market will resolve to "No".
Resolver
0x65070BE91...Anthropic's Claude family, including the recently released Claude Fable 5 and Fable 5.1 variants, currently leads LMSYS-style Overall Arena leaderboards with Elo scores near 1506 as of mid-September 2026, reflecting strong human preference in blind evaluations across writing, reasoning, and coding tasks. OpenAI's GPT-6 Astra and Meta's Muse Spark 1.2/1.3 series have narrowed gaps through iterative improvements in large language model capabilities, while open-weight contenders like Qwen3.8 Max and Kimi K3 add competitive pressure. Trader sentiment hinges on whether frontier labs can deliver further gains via scaling, post-training, or new architectures before year-end, with historical patterns showing 10-30 point monthly advances amid tight races. Key catalysts include anticipated Q4 releases, developer conferences, and any major capability demonstrations that could push the top score higher.
Resumo experimental gerado por IA com dados do Polymarket. Isto não é aconselhamento de trading e não tem qualquer papel na resolução deste mercado. · Atualizado


Cuidado com os links externos.
Cuidado com os links externos.
Frequently Asked Questions