July 16, 2026 · update
After the semis: the cheap models are on top
Two matches played. Spain beat France 2-0. Argentina beat England 2-1. The locked predictions get scored. So far, 'fast' beats 'deep'.

On 12 July we locked eight prediction boards — four models, fast tier and deep tier each — before either semi-final. Same rules: web research allowed, no betting odds.
Two games down. One to go. Here’s the score after Spain 2–0 France and Argentina 2–1 England.
Note: The main takeaway here is almost certainly not that the FAST models are better at predicting, but rather that the sport is too random (especially when matches are close) that "reasoning" about it may not be all that helpful.
Actual results (so far)
France vs Spain
England vs Argentina
Spain vs Argentina
Current points after two matches
+1 for correct winner, +2 for exact score (max 6 so far).
Only two boards correctly called both semi-finalists and can still score on the final: Claude fast (Spain vs Argentina, predicted Argentina 1–0) and Grok fast (Spain vs Argentina, predicted Spain 2–1). All deep boards and both ChatGPT boards went through France and are eliminated from final points.
What it looks like right now
The "think harder" versions mostly moved toward France in the semis (or kept France when the fast version didn’t). The fast versions of Claude and Grok landed on the actual Spain / Argentina path. On the matches played, the fast tier is strictly ahead or tied — never behind — its deep counterpart for every model.
This is, of course, two matches. The final is still to come, and only two boards can still score on it. A single 3-point swing on 19 July can reorder everything.
But the early pattern is the opposite of the comforting story: extra reasoning budget did not improve the football forecast. It made some models worse.
We’ll run the full table, final scores, and the Δ analysis after the whistle on Sunday.
The three games
France vs Spain
Dallas Stadium
England vs Argentina
Atlanta Stadium
SF winners
NY / NJ Stadium
Third place (18 Jul) is not scored. Finalists must match each model's own semi winners.
Board A
FAST — cheap / fast model
Fastest tier (Mini Light, Flash, Haiku-class, etc.). Tools and non-odds web OK.
| Model | SF1France–Spain | SF2Eng–Arg | FinalScore | Champ |
|---|---|---|---|---|
| ChatGPT | France 2–1 | England 2–1 | France 2–1 · Eng | France |
| Claude | Spain 2–1 | Argentina 2–1 | Argentina 1–0 · Spain | Argentina |
| Gemini | France 2–1 | Argentina 2–1 | Argentina 2–1 · Fra | Argentina |
| Grok | Spain 2–1 | Argentina 1–0 | Spain 2–1 · Arg | Spain |
Board B
DEEP — max reasoning
Highest-reasoning / “think hard” tier. Same info rules: research OK, no betting odds.
| Model | SF1France–Spain | SF2Eng–Arg | FinalScore | Champ | vs FAST |
|---|---|---|---|---|---|
| ChatGPT | France 2–1 | England 2–1 | France 2–1 · Eng | France | Same |
| Claude | France 2–1 | England 2–1 | France 2–1 · Eng | France | Changed |
| Gemini | France 2–1 | Argentina 2–1 | France 2–1 · Arg | France | Changed |
| Grok | France 2–1 | England 2–1 | France 1–0 · Eng | France | Changed |
Locked 12 Jul · before either semi
Cheap models disagree. Heavy models converge on France.
- ChatGPT — DEEP = FAST (identical board).
- Claude — FAST≠DEEP (Argentina path → France path).
- Gemini — FAST≠DEEP (same semis; final flips Argentina → France).
- Grok — FAST≠DEEP (Spain path → France path).
Settings
What “fast” and “deep” meant
ChatGPT
- Fast
- GPT-5.4 Mini Light · web OK, no odds
- Deep
- GPT-5.6 Sol Ultra · high reasoning + web
Claude
- Fast
- Haiku 4.5 · low effort
- Deep
- Fable 5 Max · web + extended thinking
Gemini
- Fast
- 3.1 Pro (Low) · Planning Mode
- Deep
- 3.5 Flash (High) · Planning Mode
Grok
- Fast
- Single-pass snap · no tools
- Deep
- Deliberate second pass · no odds
After 19 Jul
How we score
Correct winner
01+1 per match (SF1, SF2, Final). ET counts; pens = that side wins.
Exact score
02+2 bonus per match (reg + ET score; ignore pen digits).
Then compare
03FAST board · DEEP board · Δ (DEEP − FAST) · mean FAST vs mean DEEP.
The original locked boards (for reference)
The component above is the one published on the 12th. Nothing has been edited. The only new information is the two actual results and the points table above.
Related: the original setup post Do LLMs predict football better when they think harder? — and the API cost calculator if you want to price the difference between a fast pass and a deep one.