OpenAI’s GPT-5.4 Pro and related variants currently post 41-44% accuracy on Humanity’s Last Exam, the 2,500-question expert-level benchmark spanning mathematics, sciences, and specialized domains, placing them behind Claude Fable 5 at 53% and near Gemini 3.1 previews around 45%. With resolution in twelve days and no confirmed major model release or undisclosed optimization from OpenAI on the horizon, trader consensus assigns near-certainty to thresholds at or below 40% while viewing jumps above 50% as improbable absent sudden capability gains. Stable leaderboard trends, competitive pressure from Anthropic and Google DeepMind, and the benchmark’s resistance to rapid saturation reinforce expectations that existing frontier large language model performance will hold through June 30.
Experimental AI-generated summary referencing Polymarket data. This is not trading advice and plays no role in how this market resolves. · Updated$28,866 Vol.
35%+
Yes
40%+
Yes
50%+
No
$28,866 Vol.
35%+
Yes
40%+
Yes
50%+
No
The resolution source will be the official Humanity’s Last Exam leaderboard https://scale.com/leaderboard/humanitys_last_exam.
Market Opened: Jan 30, 2026, 12:00 AM ET
Resolver
0x65070BE91...Outcome proposed: Yes
No dispute
Final outcome: Yes
The resolution source will be the official Humanity’s Last Exam leaderboard https://scale.com/leaderboard/humanitys_last_exam.
Resolver
0x65070BE91...Outcome proposed: Yes
No dispute
Final outcome: Yes
OpenAI’s GPT-5.4 Pro and related variants currently post 41-44% accuracy on Humanity’s Last Exam, the 2,500-question expert-level benchmark spanning mathematics, sciences, and specialized domains, placing them behind Claude Fable 5 at 53% and near Gemini 3.1 previews around 45%. With resolution in twelve days and no confirmed major model release or undisclosed optimization from OpenAI on the horizon, trader consensus assigns near-certainty to thresholds at or below 40% while viewing jumps above 50% as improbable absent sudden capability gains. Stable leaderboard trends, competitive pressure from Anthropic and Google DeepMind, and the benchmark’s resistance to rapid saturation reinforce expectations that existing frontier large language model performance will hold through June 30.
Experimental AI-generated summary referencing Polymarket data. This is not trading advice and plays no role in how this market resolves. · Updated



Beware of external links.
Beware of external links.
Frequently Asked Questions