Recent releases of frontier large language models have accelerated progress on the MathArena benchmark, which evaluates uncontaminated math competition problems including AIME, HMMT, and proof-based contests. OpenAI's GPT-6 Astra, added in early September 2026, leads with expected performance near 90.7 percent across recent evaluations, ahead of Anthropic's Claude Opus 5 and Fable 5.1 series at roughly 70-72 percent. Moonshot AI's Kimi models and Alibaba's Qwen variants also post strong results on weighted math metrics, reflecting gains in reasoning chains and generalization. With multiple labs iterating rapidly through the fall, trader sentiment hinges on whether incremental improvements or a major release push any model past the market's implied threshold by year-end.
Experimentelle KI-generierte Zusammenfassung mit Polymarket-Daten. Dies ist keine Handelsberatung und spielt keine Rolle bei der Auflösung dieses Marktes. · Aktualisiert$119,848 Vol.
1575
74%
1600
33%
$119,848 Vol.
1575
74%
1600
33%
Results from the "Score" column under the "Text Arena | Math" Leaderboard tab at https://arena.ai/leaderboard/text/math-no-style-control with style control off will be used to resolve this market.
The resolution source for this market is the Chatbot Arena LLM Leaderboard found at arena.ai/leaderboard/text. If this resolution source is unavailable at check time, this market will remain open until the leaderboard comes back online and will resolve based on the first check after it becomes available. If permanently unavailable, this market will resolve to "No".
Markt eröffnet: Apr 2, 2026, 6:07 PM ET
Abwickler
0x65070BE91...Results from the "Score" column under the "Text Arena | Math" Leaderboard tab at https://arena.ai/leaderboard/text/math-no-style-control with style control off will be used to resolve this market.
The resolution source for this market is the Chatbot Arena LLM Leaderboard found at arena.ai/leaderboard/text. If this resolution source is unavailable at check time, this market will remain open until the leaderboard comes back online and will resolve based on the first check after it becomes available. If permanently unavailable, this market will resolve to "No".
Abwickler
0x65070BE91...Recent releases of frontier large language models have accelerated progress on the MathArena benchmark, which evaluates uncontaminated math competition problems including AIME, HMMT, and proof-based contests. OpenAI's GPT-6 Astra, added in early September 2026, leads with expected performance near 90.7 percent across recent evaluations, ahead of Anthropic's Claude Opus 5 and Fable 5.1 series at roughly 70-72 percent. Moonshot AI's Kimi models and Alibaba's Qwen variants also post strong results on weighted math metrics, reflecting gains in reasoning chains and generalization. With multiple labs iterating rapidly through the fall, trader sentiment hinges on whether incremental improvements or a major release push any model past the market's implied threshold by year-end.
Experimentelle KI-generierte Zusammenfassung mit Polymarket-Daten. Dies ist keine Handelsberatung und spielt keine Rolle bei der Auflösung dieses Marktes. · Aktualisiert



Vorsicht bei externen Links.
Vorsicht bei externen Links.
Häufig gestellte Fragen