Recent frontier large language model releases have driven strong trader consensus on MathArena benchmarks, where models like OpenAI's GPT-6 Astra and Anthropic's Claude Fable 5.1 currently lead with expected performance scores in the low-to-mid 80s percent range across uncontaminated math competitions. Iterative reasoning advances, tool use, and scaling have pushed top systems past 95 percent on many AIME and USAMO problems, though harder research-level and proof-based tasks remain below 75 percent for most entries. With three months left in 2026, labs are expected to ship further updates or variants ahead of year-end deadlines, potentially lifting the leader above key resolution thresholds. Open-weight challengers from Moonshot and Alibaba trail but could narrow gaps if efficiency gains accelerate adoption. Market-implied odds reflect this rapid capability trajectory tempered by typical release slippage risks.
Resumen experimental generado por IA con datos de Polymarket. Esto no es asesoramiento de trading y no influye en cómo se resuelve este mercado. · Actualizado$130,423 Vol.
1575
53%
1600
8%
$130,423 Vol.
1575
53%
1600
8%
Results from the "Score" column under the "Text Arena | Math" Leaderboard tab at https://arena.ai/leaderboard/text/math-no-style-control with style control off will be used to resolve this market.
The resolution source for this market is the Chatbot Arena LLM Leaderboard found at arena.ai/leaderboard/text. If this resolution source is unavailable at check time, this market will remain open until the leaderboard comes back online and will resolve based on the first check after it becomes available. If permanently unavailable, this market will resolve to "No".
Mercado abierto: Apr 2, 2026, 6:07 PM ET
Resolver
0x65070BE91...Results from the "Score" column under the "Text Arena | Math" Leaderboard tab at https://arena.ai/leaderboard/text/math-no-style-control with style control off will be used to resolve this market.
The resolution source for this market is the Chatbot Arena LLM Leaderboard found at arena.ai/leaderboard/text. If this resolution source is unavailable at check time, this market will remain open until the leaderboard comes back online and will resolve based on the first check after it becomes available. If permanently unavailable, this market will resolve to "No".
Resolver
0x65070BE91...Recent frontier large language model releases have driven strong trader consensus on MathArena benchmarks, where models like OpenAI's GPT-6 Astra and Anthropic's Claude Fable 5.1 currently lead with expected performance scores in the low-to-mid 80s percent range across uncontaminated math competitions. Iterative reasoning advances, tool use, and scaling have pushed top systems past 95 percent on many AIME and USAMO problems, though harder research-level and proof-based tasks remain below 75 percent for most entries. With three months left in 2026, labs are expected to ship further updates or variants ahead of year-end deadlines, potentially lifting the leader above key resolution thresholds. Open-weight challengers from Moonshot and Alibaba trail but could narrow gaps if efficiency gains accelerate adoption. Market-implied odds reflect this rapid capability trajectory tempered by typical release slippage risks.
Resumen experimental generado por IA con datos de Polymarket. Esto no es asesoramiento de trading y no influye en cómo se resuelve este mercado. · Actualizado



Cuidado con los enlaces externos.
Cuidado con los enlaces externos.
Preguntas frecuentes