I have been testing GPT-5.6 since May 27. For the first two weeks, it was the most impressive model I had ever used. Goal mode plus this model is pure sorcery. It built me a voxel-based Manhattan simulation with a working subway, and a Teardown-style destruction game, largely on its own, over runs that lasted days.
Then Claude Fable came out, and I stopped using GPT-5.6 almost overnight, because Fable is that much better for my tasks.
TL;DR
- For two weeks, the best model I'd ever used. Goal mode is a genuine step change in autonomy.
- Then Fable shipped — and despite similar benchmark scores, the gap on ambitious real work is not close.
- GPT-5.6 still wins on security work, usage limits, and the Codex agentic workflow.
Together with · Your name here
This slot reaches the people who read model reviews the day they drop
One primary sponsor per issue, written natively, clearly labeled. See the details on the advertise page.
Sponsor SBIH →The Good
Goal mode is the headline. You give the model an objective with completion criteria, and it works toward it autonomously — for days if it needs to. It's also notably strong on security work, and the usage limits are generous compared to everyone else at this tier.
Goal Mode Is Pure Sorcery
Set an objective, define what "done" means, and walk away. The runs that produced the projects below lasted days, largely unattended. There are a few techniques that reliably improve results — the numbered list in the full workflow section covers them.
It Built Manhattan
Two projects tell the story: a functioning 3D voxel recreation of Manhattan with an operational subway, and a Teardown-style destruction game with physics that hold up. I steered occasionally; the model did the work.
Then Fable Came Out
The same tasks, run head-to-head, weren't close. Fable produced superior results with minimal direction — particularly on design and creative work, where GPT-5.6 needs constant steering. Benchmarks say these models are peers. My work says otherwise.
Big Model Smell
My theory: GPT-5.6 is a smaller model with an enormous amount of reinforcement learning on top. It's exceptional at the patterns it was trained into, and it falters on genuinely novel tasks — exactly where Fable generalizes.
Where GPT-5.6 Still Wins
- Security audits — still the best model I've used for this.
- Usage limits — materially more generous than the competition.
- Codex — still the best agentic workflow interface, full stop.
When to Use What
Fable first, when you can get it. GPT-5.6 for security review and cost-conscious implementation work, and whenever you want the Codex harness specifically.
Final Thoughts
GPT-5.6 is an amazing model. I hope OpenAI's next one makes me switch back. They have done it to me before.
The next review lands the day the model drops
Reviews, guides, and the roundup. Free. No filler.