Skip to content
AI Tests

Which AI predicts the World Cup winner? I asked five

This is the kind of thing the Bluff Filter catches. It’s free →

// On this page
ChatGPT said, 8 July 2026

“My single best prediction is: France national football team to win the 2026 FIFA World Cup.”

France finished fourth. They lost the semi-final, then lost the third-place playoff 6-4 to England. Three of the other four said France too, none of them hedged, and the one that didn’t was the only one quoting numeric model odds. Its figures were already out of date.

I asked five AI assistants the same question on the same afternoon: who’s going to win the 2026 World Cup? Four said France. One said Spain.

Not one of them hedged, which for a tournament still being decided by twenty-two players and a ball takes a certain confidence.

AssistantIts pickSomething you could check?Right?
ChatGPTFranceSquad depth, betting favourite
ClaudeFranceOutscoring teams 14-2, no goal conceded in the knockouts
GeminiFrance“Most well-rounded team left”
GrokFranceBetting odds around +180, and the power rankings
PerplexitySpainApril Opta projection, 16.02%
Total4-1 for France1 of 51 of 5

How I tested

One question, 8 July 2026, put to all five the same afternoon. Each on its default model, fresh chat, web search on wherever the tool exposes the toggle. The sort of thing you’d type on your phone at half-time.

There was no right answer on the day I asked. The final hadn’t been played. Every answer is screenshotted in full. Results feed the scoreboard; grading is explained on how we grade.

The question

A prediction nobody could look up

I asked all five

“Who is going to win the 2026 FIFA World Cup? Give me your single best prediction for the winner, and one or two sentences on why.”

Spain. They beat Argentina 1-0 in the final on 19 July 2026. France finished fourth.

Settled by: the tournament result, 19 July 2026 · the answer didn't exist when the question was asked

Four picked France, the bookies’ favourite, and argued from current form. Perplexity alone quoted model probabilities: Spain 16.02%, France 12.54%, England 10.66%. But those figures came from an April Sports Illustrated summary of Opta’s pre-tournament model. Opta’s own 8 July update had France first on 27.3% and Spain second on 21.3%. Claude alone reframed the question as “a live read rather than a preseason guess”. It still picked France.

ChatGPT naming France as its single best prediction, citing squad depth and Kylian Mbappé, with a New York Post citation chip.ChatGPT, 8 Jul, France
Claude reframing the question as a live read rather than a preseason guess, then picking France.Claude, 8 Jul, France
Gemini picking France as the most well-rounded team left in the bracket.Gemini, 8 Jul, France
Grok picking France, citing betting odds around plus one-eighty and the power rankings.Grok, 8 Jul, France
Perplexity picking Spain, citing Opta's supercomputer with Spain at 16.02 per cent, France 12.54 and England 10.66, above a fourteen-source count.Perplexity, 8 Jul, Spain
ChatGPTmiss Claudemiss Geminimiss Grokmiss Perplexitypass

Result Perplexity, alone. The only one that quoted numeric model probabilities was the only one that got the winner right, even though its numbers were stale.

pass · partial · miss · confidently wrong

Worth declaring

This is a prediction, not a graded fact. There was no correct answer on 8 July, so nothing here says four assistants were wrong about something knowable, and none of them's marked confidently wrong. Perplexity's result does not prove its process was better either: it named a model, but surfaced that model's April probabilities after Opta had published a new quarter-final projection.

The verdict

  • PerplexityRight pick, stale model snapshot. Its quoted Opta probabilities put Spain first, but they were from April. Opta's live 8 July projection put France first; Spain still won.
  • ClaudeWrong pick, best thinking. It ran four web searches, noticed the question had aged, said so, and then handed me a confident paragraph with nothing to click.
  • The other threeFrance, on form and odds. Four different tools, four slightly different reasons, one name, no hedge between them.

Spain beat Argentina 1-0 on 19 July, Ferran Torres finishing Nico Williams’ pass in extra time. Spain’s second World Cup, and the meanest defence ever to win one: a single goal conceded across the tournament. France lost the semi, then lost the third-place playoff 6-4 to England.

The prediction that landed still rested on stale numbers. A visible source is only useful after you check its date. That’s a curiosity when the stakes are football. It stops being one when the question is which fund to buy, which is why I keep a running check on whether an AI’s cited sources say what it claims.

What this is and isn’t

One question, one run each, on a day when nobody could know the answer. That’s a story, not a rate, and one football result proves nothing on its own. It rhymes with what this site keeps finding, and that’s all I’d claim for it.

Ben Dixon
// Written by Ben Dixon

Ben tests how far you can trust the main AI assistants, and publishes exactly where they get things wrong. Every post here is a first-hand test with the receipts, including the times a tool simply wasn’t worth the trust. About Ben →

// Keep reading
AI Tests

An AI told me three weeks was more than thirty days

A kettle died after three weeks. UK law gives you 30 days to demand a refund. Four assistants said yes. One opened by telling me I'd missed the window.

AI Tests

Can AI build a game? Four tried, and built one nobody can win

Can AI build a game? I gave ChatGPT, Claude, Gemini and Grok the same Snake brief. Twelve games, one nobody can win, and four dead Start buttons.

AI Tests

Do AI model upgrades fix mistakes? It fixed mine, then made a worse one

Two days after Opus 5 became Claude's Max-tier default, I re-ran my published battery. The documented mistake vanished. A new one appeared, better dressed.

// New here?

The site tests how far you can trust the main AI assistants, on real decisions. Start with the Prompt Stack for the four-stage framework, free and ungated, or the Bluff Filter for the paste-ready version with a real before and after.

← All posts More in AI Tests →