Archive

July 16, 2026 · update

After the semis: the cheap models are on top

Two matches played. Spain beat France 2-0. Argentina beat England 2-1. The locked predictions get scored. So far, 'fast' beats 'deep'.

Knockout bracket with two teams left: Spain and Argentina advance to the final after the semis

On 12 July we locked eight prediction boards — four models, fast tier and deep tier each — before either semi-final. Same rules: web research allowed, no betting odds.

Two games down. One to go. Here’s the score after Spain 2–0 France and Argentina 2–1 England.

Note: The main takeaway here is almost certainly not that the FAST models are better at predicting, but rather that the sport is too random (especially when matches are close) that "reasoning" about it may not be all that helpful.

Actual results (so far)

SF114 Jul

France vs Spain

Spain 2–0
Winner: Spain
SF215 Jul

England vs Argentina

Argentina 2–1
Winner: Argentina
Final19 Jul

Spain vs Argentina

Pending

Current points after two matches

+1 for correct winner, +2 for exact score (max 6 so far).

FAST
2/2 matches
Claude (fast)4
Spain winner (+1) + Argentina exact (+3)
Gemini (fast)3
Argentina exact (+3)
Grok (fast)2
Both winners correct (+1 each)
ChatGPT (fast)0
DEEP
2/2 matches
Gemini (deep)3
Argentina exact (+3)
ChatGPT (deep)0
Claude (deep)0
Grok (deep)0

Only two boards correctly called both semi-finalists and can still score on the final: Claude fast (Spain vs Argentina, predicted Argentina 1–0) and Grok fast (Spain vs Argentina, predicted Spain 2–1). All deep boards and both ChatGPT boards went through France and are eliminated from final points.

What it looks like right now

The "think harder" versions mostly moved toward France in the semis (or kept France when the fast version didn’t). The fast versions of Claude and Grok landed on the actual Spain / Argentina path. On the matches played, the fast tier is strictly ahead or tied — never behind — its deep counterpart for every model.

This is, of course, two matches. The final is still to come, and only two boards can still score on it. A single 3-point swing on 19 July can reorder everything.

But the early pattern is the opposite of the comforting story: extra reasoning budget did not improve the football forecast. It made some models worse.

We’ll run the full table, final scores, and the Δ analysis after the whistle on Sunday.

The three games

SF1Tue 14 Jul

France vs Spain

Dallas Stadium

SF2Wed 15 Jul

England vs Argentina

Atlanta Stadium

FinalSun 19 Jul

SF winners

NY / NJ Stadium

Third place (18 Jul) is not scored. Finalists must match each model's own semi winners.

Board A

FAST — cheap / fast model

Fastest tier (Mini Light, Flash, Haiku-class, etc.). Tools and non-odds web OK.

ModelSF1France–SpainSF2Eng–ArgFinalScoreChamp
ChatGPT
France
2–1
England
2–1
France
2–1 · Eng
France
Claude
Spain
2–1
Argentina
2–1
Argentina
1–0 · Spain
Argentina
Gemini
France
2–1
Argentina
2–1
Argentina
2–1 · Fra
Argentina
Grok
Spain
2–1
Argentina
1–0
Spain
2–1 · Arg
Spain

Board B

DEEP — max reasoning

Highest-reasoning / “think hard” tier. Same info rules: research OK, no betting odds.

ModelSF1France–SpainSF2Eng–ArgFinalScoreChampvs FAST
ChatGPT
France
2–1
England
2–1
France
2–1 · Eng
FranceSame
Claude
France
2–1
England
2–1
France
2–1 · Eng
FranceChanged
Gemini
France
2–1
Argentina
2–1
France
2–1 · Arg
FranceChanged
Grok
France
2–1
England
2–1
France
1–0 · Eng
FranceChanged

Locked 12 Jul · before either semi

Cheap models disagree. Heavy models converge on France.

  • ChatGPT — DEEP = FAST (identical board).
  • Claude — FAST≠DEEP (Argentina path → France path).
  • Gemini — FAST≠DEEP (same semis; final flips Argentina → France).
  • Grok — FAST≠DEEP (Spain path → France path).
FAST champsFranceArgentinaArgentinaSpain
DEEP champsFranceFranceFranceFrance

Settings

What “fast” and “deep” meant

ChatGPT

Fast
GPT-5.4 Mini Light · web OK, no odds
Deep
GPT-5.6 Sol Ultra · high reasoning + web

Claude

Fast
Haiku 4.5 · low effort
Deep
Fable 5 Max · web + extended thinking

Gemini

Fast
3.1 Pro (Low) · Planning Mode
Deep
3.5 Flash (High) · Planning Mode

Grok

Fast
Single-pass snap · no tools
Deep
Deliberate second pass · no odds

After 19 Jul

How we score

Correct winner

01

+1 per match (SF1, SF2, Final). ET counts; pens = that side wins.

Exact score

02

+2 bonus per match (reg + ET score; ignore pen digits).

Then compare

03

FAST board · DEEP board · Δ (DEEP − FAST) · mean FAST vs mean DEEP.

The original locked boards (for reference)

The component above is the one published on the 12th. Nothing has been edited. The only new information is the two actual results and the points table above.

Related: the original setup post Do LLMs predict football better when they think harder? — and the API cost calculator if you want to price the difference between a fast pass and a deep one.