Vector Lab
VECTOR LAB

EST. 2025

WEEKLY UPDATE2026
BY ANDREW MEAD

GPT 6 initial assessment

What are the initial vibes for GPT 6, and how does it fare compared to Fable 5.1

Read in

Releases

GPT 6 Astra

GPT 6 Astra got released this week, with most people getting access to it on Friday.

Because of this, we have not had enough time to get a good sense of the model’s full capabilities yet (as we saw with Opus 5 it takes 3-4 days to learn how well the model truly performs in the real world) but it does give a preview of what to look out for as you are using it this week.

We will do a more thorough breakdown next week, but for now we will go over what we do know about the model and what to potentially look out for.

GPT 6 astra benchmarks

GPT 6’s most notable improvements come from its increased vision capabilities. It maxes out various private vision benchmarks, exceeding human baseline scores on them.

This then leads it to have very strong browser/computer use, and impressive performance in 3D tools like Blender and CAD programs. This of course is the ideal thing to post about on social media (a cool Blender render gets far more engagement than a generic benchmark chart), and does showcase how the model’s vision capabilities are by far and away the best of any model we have ever seen, but I think this may be overinflating how good we think the model is.

This is not image gen. GPT-6 designed all of these in Bricklink Studio (brick by brick) then rendered them in Blender. From dominik kundel on Twitter

We can see a few cracks potentially forming, some people have reported very low code quality (worse than any previous frontier model) which may also be an indication of poor writing abilities in general.

It has also been found that the model underperforms GPT 5.6 and Fable when given multiple attempts at a problem, which could be a sign of brittleness and lack of diversity in the model, which would impact things like auto research.

This is not to say that the model is bad, just that further investigation is needed, and that the model is not some silver bullet that has fixed every issue that we have had with LLMs.

Then there is the issue of price. Similar to previous major releases from OpenAI, GPT 6 comes with a price hike from $30 to $50 per million output tokens, but claims to be about equal in price due to a decrease in token usage since the model is much smarter and efficient.

Model$ per million (input)$ per million (output)Tokens per second
GPT 6 Astra$10$5035
GPT 5.6 Sol$5$3045
Claude Fable 5.1$10$5040
GLM 5.3$1.40$4.4047
Grok 4.6$2$656

The issue is that for easy tasks, GPT 5.6 Sol was already as efficient as you could be for them, and there are no big token savings to be had, which means for those easy tasks you will be paying more. On the Artificial Analysis benchmark suite, they find that GPT 6 with low reasoning is more expensive that GPT 5.6 Sol high, even with a 70% increase in token efficiency.

GPT 6 vs GPT 5.6 costs on AA

Because of this, I would expect Codex limits to drain faster than before, and if you are using it via the API be sure to monitor its costs since it will be notably higher than its predecessor.

I will continue to use the model this week and gather feedback from the community, and will have a more thorough writeup for next week, so be sure to subscribe to the newsletter so that you don’t miss out.

Fable 5.1

Anthropic released an update to Fable and Mythos a few days before the release of GPT 6, to much less fanfare. For those that are unaware, Fable and Mythos are the same model, with Fable having extra safety training and guardrails added to it so that the general public can use it without causing too much havoc. Mythos is only available to select organizations who are participating in Anthropic’s Project Glasswing, which is focused on cybersecurity.

Fable 5.1 benchmarks

They did not bother making a separate benchmark image, this is just a screenshot from the system card

Fable 5.1 is a marginal upgrade to its predecessor, with the main notable difference being the drastic reduction in jargon that the model uses to communicate, making it feel much more natural to talk and work with.

The only other notable changes come with token use and pricing. The model uses more tokens, which is caused by its reluctance to use parallel tool calls, and desire to completely rewrite files when performing code edits, which seem like a training issue from Anthropic, as I can’t imagine this the best way to be working.

This increase in token use is offset by a 75% reduction in cached input token costs (which are the majority of the tokens that you use in agentic applications). These changes somewhat offset each other, ending up with a 20% increase in costs according to Artificial Analysis.

You should also note that when using Fable 5.1 on the low reasoning setting, it is more likely to use its own internal knowledge instead of using its web search tool, which could cause it to hallucinate more or tell you out of date information.

Fable 5.1 vs GPT 6 making a villa in Blender from hesamation on Twitter

Finish

I hope you enjoyed the news this week. If you want to get the news every week, be sure to join our mailing list below.

Arch in the desert at dusk

From Tatiana Tsiguleva on Twitter

Stay Updated

Subscribe to get the latest AI news in your inbox every week!

← BACK TO NEWS