OpenAI’s GPT-5.4 Pro and related variants currently post 41-44% accuracy on Humanity’s Last Exam, the 2,500-question expert-level benchmark spanning mathematics, sciences, and specialized domains, placing them behind Claude Fable 5 at 53% and near Gemini 3.1 previews around 45%. With resolution in twelve days and no confirmed major model release or undisclosed optimization from OpenAI on the horizon, trader consensus assigns near-certainty to thresholds at or below 40% while viewing jumps above 50% as improbable absent sudden capability gains. Stable leaderboard trends, competitive pressure from Anthropic and Google DeepMind, and the benchmark’s resistance to rapid saturation reinforce expectations that existing frontier large language model performance will hold through June 30.
Eksperymentalne podsumowanie AI odwołujące się do danych Polymarket. To nie jest porada handlowa i nie ma wpływu na rozstrzyganie tego rynku. · ZaktualizowanoOpenAI GPT score on Humanity’s Last Exam by June 30?
$28,866 Wol.
35%+
Yes
40%+
Yes
50%+
No
$28,866 Wol.
35%+
Yes
40%+
Yes
50%+
No
The resolution source will be the official Humanity’s Last Exam leaderboard https://scale.com/leaderboard/humanitys_last_exam.
Rynek otwarty: Jan 30, 2026, 12:00 AM ET
Rozstrzygający
0x65070be91...Wynik zaproponowany: Yes
Brak sporu
Ostateczny wynik: Yes
The resolution source will be the official Humanity’s Last Exam leaderboard https://scale.com/leaderboard/humanitys_last_exam.
Rozstrzygający
0x65070be91...Wynik zaproponowany: Yes
Brak sporu
Ostateczny wynik: Yes
OpenAI’s GPT-5.4 Pro and related variants currently post 41-44% accuracy on Humanity’s Last Exam, the 2,500-question expert-level benchmark spanning mathematics, sciences, and specialized domains, placing them behind Claude Fable 5 at 53% and near Gemini 3.1 previews around 45%. With resolution in twelve days and no confirmed major model release or undisclosed optimization from OpenAI on the horizon, trader consensus assigns near-certainty to thresholds at or below 40% while viewing jumps above 50% as improbable absent sudden capability gains. Stable leaderboard trends, competitive pressure from Anthropic and Google DeepMind, and the benchmark’s resistance to rapid saturation reinforce expectations that existing frontier large language model performance will hold through June 30.
Eksperymentalne podsumowanie AI odwołujące się do danych Polymarket. To nie jest porada handlowa i nie ma wpływu na rozstrzyganie tego rynku. · Zaktualizowano



Uważaj na linki zewnętrzne.
Uważaj na linki zewnętrzne.
Często zadawane pytania