Anthropic currently holds the edge in most independent benchmarks following its early September release of Claude Fable 5.1 and Mythos 5.1, which top composites like BenchLM.ai at 84.74 and lead in agentic coding, reasoning, and knowledge work tasks ahead of OpenAI’s GPT-6 Astra. OpenAI’s September 3 launch of Astra keeps it in close contention at #2 across many evaluations, while Google’s Gemini 3.8 Flash, Meta’s Muse Spark 1.3, and models from xAI and Alibaba trail in the top tier but excel in specific areas like multimodality or cost efficiency. Frequent frontier updates, benchmark rebaselining, and internal safety discussions among labs create ongoing volatility through year-end, with any additional Q4 releases or capability demonstrations likely to shift relative positioning among the leading developers.
Tóm tắt AI thử nghiệm tham chiếu dữ liệu Polymarket. Đây không phải tư vấn giao dịch và không ảnh hưởng đến cách thị trường này được giải quyết. · Cập nhật$158,069 KL.
27%
OpenAI
25%
xAI
13%
Z.ai
11%
Meta
10%
Moonshot
8%
Alibaba
7%
Baidu
7%
DeepSeek
6%
Microsoft
5%
Amazon
3%
Mistral
3%
Meituan
3%
ByteDance
1%
$158,069 KL.
27%
OpenAI
25%
xAI
13%
Z.ai
11%
Meta
10%
Moonshot
8%
Alibaba
7%
Baidu
7%
DeepSeek
6%
Microsoft
5%
Amazon
3%
Mistral
3%
Meituan
3%
ByteDance
1%
Results from the "Rank" column under the "Text Arena | Overall" Leaderboard tab at https://lmarena.ai/leaderboard/text with style control off will be used to resolve this market.
If a listed model ties for #1 Arena rank, it will suffice to resolve this market to "Yes."
The resolution source for this market is the Chatbot Arena LLM Leaderboard found at https://lmarena.ai/. If this resolution source is unavailable at check time, this market will remain open until the leaderboard comes back online and will resolve based on the first check after it becomes available. If it becomes permanently unavailable, this market will resolve based on another resolution source.
Thị trường mở: Apr 30, 2026, 3:22 PM ET
Người giải quyết
0x65070BE91...Results from the "Rank" column under the "Text Arena | Overall" Leaderboard tab at https://lmarena.ai/leaderboard/text with style control off will be used to resolve this market.
If a listed model ties for #1 Arena rank, it will suffice to resolve this market to "Yes."
The resolution source for this market is the Chatbot Arena LLM Leaderboard found at https://lmarena.ai/. If this resolution source is unavailable at check time, this market will remain open until the leaderboard comes back online and will resolve based on the first check after it becomes available. If it becomes permanently unavailable, this market will resolve based on another resolution source.
Người giải quyết
0x65070BE91...Anthropic currently holds the edge in most independent benchmarks following its early September release of Claude Fable 5.1 and Mythos 5.1, which top composites like BenchLM.ai at 84.74 and lead in agentic coding, reasoning, and knowledge work tasks ahead of OpenAI’s GPT-6 Astra. OpenAI’s September 3 launch of Astra keeps it in close contention at #2 across many evaluations, while Google’s Gemini 3.8 Flash, Meta’s Muse Spark 1.3, and models from xAI and Alibaba trail in the top tier but excel in specific areas like multimodality or cost efficiency. Frequent frontier updates, benchmark rebaselining, and internal safety discussions among labs create ongoing volatility through year-end, with any additional Q4 releases or capability demonstrations likely to shift relative positioning among the leading developers.
Tóm tắt AI thử nghiệm tham chiếu dữ liệu Polymarket. Đây không phải tư vấn giao dịch và không ảnh hưởng đến cách thị trường này được giải quyết. · Cập nhật



Cẩn thận với liên kết bên ngoài.
Cẩn thận với liên kết bên ngoài.
Câu hỏi thường gặp