Recent model releases have intensified competition on Code Arena's WebDev leaderboard, a crowdsourced Elo system evaluating agentic coding through blind community votes on real frontend tasks. OpenAI's GPT-6 Astra launched September 3 and quickly claimed the top spot near 1,800 points, narrowly ahead of Anthropic's Claude Fable 5.1-max. Other frontier systems from Anthropic, OpenAI, and labs like Alibaba and DeepSeek continue closing gaps on multi-step reasoning and tool use. Traditional benchmarks such as SWE-bench have largely saturated above 90 percent, shifting trader focus to live arenas that better reflect iterative development. Additional high-parameter releases and capability jumps expected before year-end could push scores higher, though gains depend on sustained progress in long-horizon agent performance.
基于Polymarket数据的AI实验性摘要。这不是交易建议,也不影响该市场的结算方式。 · 更新于$213,543 交易量
1560
32%
1580
19%
1600
5%
$213,543 交易量
1560
32%
1580
19%
1600
5%
Results from the "Score" column under the "Text Arena | Coding" Leaderboard tab at https://arena.ai/leaderboard/text/coding-no-style-control with style control off will be used to resolve this market.
The resolution source for this market is the Chatbot Arena LLM Leaderboard found at arena.ai/leaderboard/text. If this resolution source is unavailable at check time, this market will remain open until the leaderboard comes back online and will resolve based on the first check after it becomes available. If permanently unavailable, this market will resolve to "No".
市场开放时间: Apr 2, 2026, 6:09 PM ET
Results from the "Score" column under the "Text Arena | Coding" Leaderboard tab at https://arena.ai/leaderboard/text/coding-no-style-control with style control off will be used to resolve this market.
The resolution source for this market is the Chatbot Arena LLM Leaderboard found at arena.ai/leaderboard/text. If this resolution source is unavailable at check time, this market will remain open until the leaderboard comes back online and will resolve based on the first check after it becomes available. If permanently unavailable, this market will resolve to "No".
Recent model releases have intensified competition on Code Arena's WebDev leaderboard, a crowdsourced Elo system evaluating agentic coding through blind community votes on real frontend tasks. OpenAI's GPT-6 Astra launched September 3 and quickly claimed the top spot near 1,800 points, narrowly ahead of Anthropic's Claude Fable 5.1-max. Other frontier systems from Anthropic, OpenAI, and labs like Alibaba and DeepSeek continue closing gaps on multi-step reasoning and tool use. Traditional benchmarks such as SWE-bench have largely saturated above 90 percent, shifting trader focus to live arenas that better reflect iterative development. Additional high-parameter releases and capability jumps expected before year-end could push scores higher, though gains depend on sustained progress in long-horizon agent performance.
基于Polymarket数据的AI实验性摘要。这不是交易建议,也不影响该市场的结算方式。 · 更新于



警惕外部链接哦。
警惕外部链接哦。
常见问题