Vector Lab
VECTOR LAB

EST. 2025

WEEKLY UPDATE2026
BY ANDREW MEAD

Opus 5.5 vs GPT 6.1 Sol

Opus and Sonnet 5.5, GPT 6.1 Sol and 6 Luna, OpenAI Dev Day Recap, and more!

Read in

Note

Sorry there was no AI news last week, I was out in the woods camping so was completely unaware of all the new models and news that came out during the week. That coupled with all of the new models released this week will make for the biggest AI news yet. I am once again going away on a trip this weekend, so I am writing this on Thursday afternoon, hopefully nothing else important happens after today.

News

OpenAI Dev Day

OpenAI had their yearly Dev Day this week where they announced a number of things, here are some of the highlights.

Dots, an AI personal assistant that is meant to be a competitor to OpenClaw, Meta Muse, and GrokBot. Seems fairly similar capability-wise to the competition, which makes sense since OpenAI owns OpenClaw already. It works directly with your subscription, and for the first few weeks has effectively infinite usage.

Decisions API: OpenAI’s Jev competitor. It uses GPT-6 Luna under the hood to perform zero shot classification, making it more expensive than Jev (about 2x the price for input tokens, and you have to pay for output tokens as well) but with the benefit of being able to handle multimodal inputs. They did not release any benchmarks with it (red flag), which coupled with the higher costs, makes it a pass for me right now.

A new subscription plan for $500 a month. It mostly just gets you access to more usage (no need to juggle as many accounts), with the only notable feature being access to UltraFast, which gives 5-10x faster inference speeds for Astra (300 tokens per second), but at 6x higher cost/ rate limit usage.

Sign in with ChatGPT. Developers can now add a sign in with ChatGPT option to their sites, and when a user does so they can use their subscription on that site, allowing them to bring their subscription wherever they go. OpenAI seems to be taking the opposite stance from Anthropic, who are closing down the ecosystem around their models and where they can use them.

OpenAI also announced that enterprise users can use 3rd party open source models through their partner BaseTen. This means that enterprises will have access to models like GLM 5.3 and Kimi K3 through their OpenAI deals. This is also a vast departure from Anthropic, who just this week called GLM 5.3 a dangerous model due to its lack of guardrails.

All of these announcements will probably have less effect on you than this final one that they were hiding during Dev Day, which is that rate limits (in terms of API dollars) on the subscription plan are being halved. They say this is due to the API prices of their models going down dramatically. For instance Sol and Luna saw a 50% price decrease last week, so your usage for those will feel unchanged. The main way you will notice is for Astra usage, since that has not seen any price cuts, so it will feel like you get to use it half as much.

Releases

Opus 5.5

Anthropic’s big release from 2 weeks ago was Opus 5.5.

Opus 5.5 benchmarks

The previous couple of Opus models have been a bit lackluster, with the last great Opus model being 4.6. Opus 5.5 bucks that trend, as this model appears to be the real deal now and not just a model that moves the benchmark needle.

The thing that stands out about this model is its artistic taste. There are many examples on Twitter of it using javascript to make artwork or putting together a unique and interesting music video for a song.

Video from NotinReality on Twitter

The issues with the previous Claude models was the unreliability and overzealousness on tasks. They would often skip over things you would have wanted them to do, or go and add additional features or complexity that you did not ask for. Anthropic has fixed this wide variety in outcomes, making it much more usable day to day.

They also have decreased the pricing for it by 20% compared to previous versions of Opus, and drastically reduced the number of reasoning tokens it uses, especially at lower reasoning levels (low, medium, and high), making it even cheaper than the decrease in sticker price would entail.

Model$ per million (input)$ per million (output)Tokens per second
Claude Opus 5.5$4$2068
Claude Opus 5$5$2569
Claude Sonnet 5.5$2$1084
GPT 6.1 Sol$2$1031
GPT 6 Luna$0.10$0.5035
GPT 6 Astra$10$5035
Claude Fable 5.1$10$5040
GLM 5.3$1.40$4.4047

Anthropic models have always been better at inferring the intent of your prompt vs OpenAI’s models, and they also tend to write high quality code, but they have lacked a good option to use since Fable burns through limits fast and you can only use it for up to 50% of your subscription quota.

Opus 5.5 is bringing people back to Claude from Codex and OpenAI, as people had been migrating away after so many lackluster releases from Anthropic. I haven’t had too much time with the model yet but I will certainly be dusting off my Claude Code subscription to try out this model more in the coming weeks.

New GPT models

GPT 6 Sol and Luna

We will start with the models from 2 weeks ago that I was not able to cover. GPT-6 Sol and Luna are mostly refreshes of the 5.6 models. With the GPT-6 release Terra got dropped entirely, leaving just Sol, Luna, and Astra, and many people seemed to think that GPT-6 Sol was just a rebranded Terra, but that is now irrelevant with the 6.1 Sol release which we will talk about in the moment.

Luna, Sol, and Astra

Notably, both Sol and Luna saw drastic 50% price decreases, and also token usage improvements, making them even more affordable than they were before, with Luna now being 1/100th the price of Astra.

We have not talked much about Luna outside of the initial release, but it has turned out to be one of the best price to performance models on the market right now. It is now the default model that I use in agentic systems, as it is fast, low cost, and fairly intelligent.

GPT 6.1 Sol

GPT-6 Sol’s initial poor performance got quickly released with the 6.1 version. This model feels much more like what we expected GPT-6 Sol to be, with benchmarks rivaling Astra at 1/5th the price.

DeepSWE benchmarks for the GPT 6 models

As expected the model is around Astra level, with it difficult to distinguish between the two. Compared to Opus 5.5, the lack of taste and more rigid instruction following becomes apparent, and I have seen many people who have tried both decide to go with Opus 5.5 instead of Sol 6.1 for their daily driver.

One of the big things I have noticed with it this week is its lack of speed. By making it cheaper many more people have started using it, and they have failed to scale up enough to meet demand. Even though it is fairly token efficient, running at 20 tokens per seconds means that most prompts still take 5-10 minutes to complete, whereas Opus 5.5 is running around 70 tokens per second, making it more than 3x faster.

I would still consider it a very good model, and place it right below Opus 5.5 as the best day to day model to use right now.

Quick Hits

Gemini 4 Argon

Google saw Anthropic and OpenAI releasing models and wanted to get in on the action as well, but their model is just more of the same old Gemini BS.

Gemini 4 Argon benchmarks

The model benchmarks well, but in the real world its intelligence is very jagged and unpleasant to use. From leaked conversations from within the Gemini team, by their own admission, the model is behind on coding compared to Anthropic and OpenAI’s offerings.

At $4 per million input tokens and $20 for output, it is not on the cost/intelligence frontier at all as well, making it an easy pass.

Sonnet 5.5

Sonnet 5.5 is another breath of fresh air for the Sonnet models similar to Opus 5.5.

Sonnet 5.5 benchmarks

Capability wise it is fairly similar to Opus, with higher token usage, making it about 30% cheaper in the real world at the same reasoning level as Opus.

The main way to describe the difference between Opus and Sonnet 5.5 is that Opus has wisdom that comes from being a larger model with Sonnet does not. This means that Opus will be better at planning and decision making vs Sonnet.

Given Opus’ price decrease and its better capabilities, I find it hard to recommend Sonnet unless you are really trying to maximize your Claude Code usage (I would use Opus as a planning agent and Sonnet as the implementer). For most people though, I would just recommend using Opus and saving Sonnet for your easier tasks.

Also a part of this announcement was that Haiku is going to return in the future (it hasn’t been updated for almost 2 years now) and I think Opus 5.5 and haiku 5.5 will be the best combo to use together, further making Sonnet feel irrelevant.

Nemotron 3 Diarization

Amongst all these LLM releases, I wanted to briefly mention that we have a new SOTA diarization model from Nvidia.

Diarization models split a single audio track with multiple voices into individual voice tracks, making it easy to isolate which speaker is saying what. Previous models were all rather lack luster, and usually behind some kind of pay wall, but the Nemotron 3 Diarization model is both open source and has a 25% lower error rate than the next best model on Diarization Bench.

This is great for things like podcast transcriptions, allowing you to break it into individual tracks and sort your data into who said what and when, which can be very beneficial when passing information to an LLM, since it does not have native audio listen capabilities on its own (unless you are using a Gemini model).

Finish

I hope you enjoyed the news this week. If you want to get the news every week, be sure to join our mailing list below.

VHS Tape cover looking ahh art

Code art generation by Opus 5.5 from owl on Twitter

Stay Updated

Subscribe to get the latest AI news in your inbox every week!

← BACK TO NEWS