SOMETHING BIG IS HAPPENING · by Matt Shumer New AI reviews · How-tos · The Roundup · Advertise 2026·07·21

← All writing

Review · Day oneGPT-5.6

My GPT-5.6 Review: Second Place Has Never Been This Good

OpenAI's new flagship family is the most impressive thing I'd ever used — for two weeks. Then Fable came out. Here's the full accounting.

I have been testing GPT-5.6 since May 27. For the first two weeks, it was the most impressive model I had ever used. Goal mode plus this model is pure sorcery. It built me a voxel-based Manhattan simulation with a working subway, and a Teardown-style destruction game, largely on its own, over runs that lasted days.

Then Claude Fable came out, and I stopped using GPT-5.6 almost overnight, because Fable is that much better for my tasks.

TL;DR

  • For two weeks, the best model I'd ever used. Goal mode is a genuine step change in autonomy.
  • Then Fable shipped — and despite similar benchmark scores, the gap on ambitious real work is not close.
  • GPT-5.6 still wins on security work, usage limits, and the Codex agentic workflow.

The Good

Goal mode is the headline. You give the model an objective with completion criteria, and it works toward it autonomously — for days if it needs to. It's also notably strong on security work, and the usage limits are generous compared to everyone else at this tier.

Goal Mode Is Pure Sorcery

Set an objective, define what "done" means, and walk away. The runs that produced the projects below lasted days, largely unattended. There are a few techniques that reliably improve results — the numbered list in the full workflow section covers them.

It Built Manhattan

Voxel Manhattan — aerial view
The simulation includes a working subway system. Built largely autonomously over multi-day runs.

Two projects tell the story: a functioning 3D voxel recreation of Manhattan with an operational subway, and a Teardown-style destruction game with physics that hold up. I steered occasionally; the model did the work.

Then Fable Came Out

The same tasks, run head-to-head, weren't close. Fable produced superior results with minimal direction — particularly on design and creative work, where GPT-5.6 needs constant steering. Benchmarks say these models are peers. My work says otherwise.

Big Model Smell

My theory: GPT-5.6 is a smaller model with an enormous amount of reinforcement learning on top. It's exceptional at the patterns it was trained into, and it falters on genuinely novel tasks — exactly where Fable generalizes.

Where GPT-5.6 Still Wins

  • Security audits — still the best model I've used for this.
  • Usage limits — materially more generous than the competition.
  • Codex — still the best agentic workflow interface, full stop.

When to Use What

Fable first, when you can get it. GPT-5.6 for security review and cost-conscious implementation work, and whenever you want the Codex harness specifically.

Final Thoughts

GPT-5.6 is an amazing model. I hope OpenAI's next one makes me switch back. They have done it to me before.

The next review lands the day the model drops

Reviews, guides, and the roundup. Free. No filler.