OpenAI’s GPT-5.4 Pro and related variants currently post 41-44% accuracy on Humanity’s Last Exam, the 2,500-question expert-level benchmark spanning mathematics, sciences, and specialized domains, placing them behind Claude Fable 5 at 53% and near Gemini 3.1 previews around 45%. With resolution in twelve days and no confirmed major model release or undisclosed optimization from OpenAI on the horizon, trader consensus assigns near-certainty to thresholds at or below 40% while viewing jumps above 50% as improbable absent sudden capability gains. Stable leaderboard trends, competitive pressure from Anthropic and Google DeepMind, and the benchmark’s resistance to rapid saturation reinforce expectations that existing frontier large language model performance will hold through June 30.
Experimentelle KI-generierte Zusammenfassung mit Polymarket-Daten. Dies ist keine Handelsberatung und spielt keine Rolle bei der Auflösung dieses Marktes. · Aktualisiert$28,866 Vol.
35 %+
Ja
40 %+
Ja
50 %+
Nein
$28,866 Vol.
35 %+
Ja
40 %+
Ja
50 %+
Nein
The resolution source will be the official Humanity’s Last Exam leaderboard https://scale.com/leaderboard/humanitys_last_exam.
Markt eröffnet: Jan 30, 2026, 12:00 AM ET
Abwickler
0x65070be91...Vorgeschlagenes Ergebnis: Ja
Kein Einspruch
Endgültiges Ergebnis: Ja
The resolution source will be the official Humanity’s Last Exam leaderboard https://scale.com/leaderboard/humanitys_last_exam.
Abwickler
0x65070be91...Vorgeschlagenes Ergebnis: Ja
Kein Einspruch
Endgültiges Ergebnis: Ja
OpenAI’s GPT-5.4 Pro and related variants currently post 41-44% accuracy on Humanity’s Last Exam, the 2,500-question expert-level benchmark spanning mathematics, sciences, and specialized domains, placing them behind Claude Fable 5 at 53% and near Gemini 3.1 previews around 45%. With resolution in twelve days and no confirmed major model release or undisclosed optimization from OpenAI on the horizon, trader consensus assigns near-certainty to thresholds at or below 40% while viewing jumps above 50% as improbable absent sudden capability gains. Stable leaderboard trends, competitive pressure from Anthropic and Google DeepMind, and the benchmark’s resistance to rapid saturation reinforce expectations that existing frontier large language model performance will hold through June 30.
Experimentelle KI-generierte Zusammenfassung mit Polymarket-Daten. Dies ist keine Handelsberatung und spielt keine Rolle bei der Auflösung dieses Marktes. · Aktualisiert



Vorsicht bei externen Links.
Vorsicht bei externen Links.
Häufig gestellte Fragen