OpenAI’s GPT-5.4 Pro and related variants currently post 41-44% accuracy on Humanity’s Last Exam, the 2,500-question expert-level benchmark spanning mathematics, sciences, and specialized domains, placing them behind Claude Fable 5 at 53% and near Gemini 3.1 previews around 45%. With resolution in twelve days and no confirmed major model release or undisclosed optimization from OpenAI on the horizon, trader consensus assigns near-certainty to thresholds at or below 40% while viewing jumps above 50% as improbable absent sudden capability gains. Stable leaderboard trends, competitive pressure from Anthropic and Google DeepMind, and the benchmark’s resistance to rapid saturation reinforce expectations that existing frontier large language model performance will hold through June 30.
Riepilogo sperimentale generato dall'AI con riferimento ai dati di Polymarket. Questo non è un consiglio di trading e non ha alcun ruolo nella risoluzione di questo mercato. · AggiornatoPunteggio OpenAI GPT all'ultimo esame dell'umanità entro il 30 giugno?
$28,866 Vol.
35%+
Sì
40%+
Sì
50%+
No
$28,866 Vol.
35%+
Sì
40%+
Sì
50%+
No
The resolution source will be the official Humanity’s Last Exam leaderboard https://scale.com/leaderboard/humanitys_last_exam.
Mercato aperto: Jan 30, 2026, 12:00 AM ET
Risolutore
0x65070be91...Esito proposto: Sì
Nessuna contestazione
Esito finale: Sì
The resolution source will be the official Humanity’s Last Exam leaderboard https://scale.com/leaderboard/humanitys_last_exam.
Risolutore
0x65070be91...Esito proposto: Sì
Nessuna contestazione
Esito finale: Sì
OpenAI’s GPT-5.4 Pro and related variants currently post 41-44% accuracy on Humanity’s Last Exam, the 2,500-question expert-level benchmark spanning mathematics, sciences, and specialized domains, placing them behind Claude Fable 5 at 53% and near Gemini 3.1 previews around 45%. With resolution in twelve days and no confirmed major model release or undisclosed optimization from OpenAI on the horizon, trader consensus assigns near-certainty to thresholds at or below 40% while viewing jumps above 50% as improbable absent sudden capability gains. Stable leaderboard trends, competitive pressure from Anthropic and Google DeepMind, and the benchmark’s resistance to rapid saturation reinforce expectations that existing frontier large language model performance will hold through June 30.
Riepilogo sperimentale generato dall'AI con riferimento ai dati di Polymarket. Questo non è un consiglio di trading e non ha alcun ruolo nella risoluzione di questo mercato. · Aggiornato



Fai attenzione ai link esterni.
Fai attenzione ai link esterni.
Domande frequenti