Back to Nick Test
Source

Opus 5.5 vs. GPT-6 Sol: which model won my blind taste test?

NT
Nick Test
@nick-test

I got up early to record an Opus 5.5 review. Then Anthropic and OpenAI dropped new models on the same morning, and I decided to do something I’d never done before: take the How I AI bench live. I put GPT-6 Astra, GPT-6 Sol, Claude Opus 5.5, and more through the work I actually care about: emails, PRDs, frontend prototypes, backend work, long-running agents, SVGs, and video editing. I scored the outputs without knowing which model made them, so you get to watch me make predictions, change my mind, and reveal my own very inconsistent taste. Astra won my heart. Opus 5.5 won my week. Sol still has me split. There’s a creative result I got completely wrong, an LLM judge that disagreed with me, and a return to Barbie Bench: the 3D fashion game that keeps reminding me how far we have to go. The hands are tragic. AGI has not arrived.

Appears in

Uploaded
Uploaded Sep 23, 2026
File type
POD
Queried
0

Full transcript

Showing the full transcript for this episode.

No preview text is available for this document yet.

Want to learn more?

Ask about this episode