Vector Lab
VECTOR LAB

EST. 2025

WEEKLY UPDATE2026
BY ANDREW MEAD

DeepSeek V4 Pro

Can DeepSeek compete with Fable and Sol? Is SpaceXAI becoming a frontier lab? Qwen returns to small open source models

Read in

Releases

DeepSeek V4 Pro

A few weeks ago DeepSeek released the updated version of the V4 Flash model, which included much better post training versus its predecessor, making it one of the best low cost options for day to day LLM tasks.

Their DeepSeek V4 Pro update aims to be a competitor at the top along side GPT 5.6 and Fable, and not just the value king.

Deepseek v4 pro benchmark scores

Compared to the impressive amount of intelligence that was packed into the V4 Flash model, V4 Pro is a bit of a flop. On Artificial Analysis it barely scores better than Flash, while being ~2.5x more expensive. Like Flash, it is not multimodal, something that the model itself does not like.

In the real world, it seems to do better at handling the planning for longer duration agentic tasks, but its actual coding performance is not that much different.

Further hindering the model is the pricing increases they are enacting (about 3x for Flash and 4x for Pro). This puts them squarely in competition with GPT 5.6 Luna and Terra pricing wise, once factoring in the increased number of tokens that the DeepSeek models use on average.

Deepseek pricing

Model$ per million (input)$ per million (output)Tokens per second
DeepSeek V4 Pro$1.32$3.9655
DeepSeek V4 Flash$0.44$1.3282
Grok 4.6$2$657
Claude Fable 5$10$5040
Claude Opus 4.8$5$2560
GPT 5.6 Sol$5$3045
GPT 5.6 Terra$2.50$1563
GPT 5.6 Luna$1$684
Kimi K3$3$1527
GLM 5.2$1.40$4.4047

There is a silver lining with this however. Flash was most likely distilled from an early version of the Pro model, and the relative parity it has with Pro means that DeepSeek is probably the best in the world at distilling their models down to a smaller size, and is part of why this release is so underwhelming.

If this continues long term, this means that DeepSeek will be able to compact frontier level intelligence down into smaller and cheaper models that we will be able to use and potentially run at home ourselves.

I’d skip V4 Pro right now. If you need a cheap model, look at V4 Flash, but Pro sits in the mushy middle where it is neither the cheapest nor the smartest, making it hard for me to recommend.

Grok 4.6

After a long time in the “meh” category of models, it seems like Grok is finally making a name for itself with iteration 4.6.

Grok 4.6 benchmarks

For day to day coding tasks, from what I’ve seen, it seems to be on part with GPT 5.6 Sol and Fable, and the initial vibes seem to place it above most of the Chinese models, probably around the Kimi K3 level or better. For large scale tasks and raw reasoning capability it falls behind the frontier models, so it is not a fully replacement for GPT 5.6 and Fable, but for most people this is probably not an issue.

Similar to its predecessor, many people have talked about how snappy the model feels to use when compared to the competition, especially when using the “low” reasoning setting.

Its also very impressive to see the speed that the SpaceXAI team is iterating at right now. Grok 4.5 was released one month prior, and they have already announced that Grok 4.7 will be released in month. This is fast even for the breakneck speeds in the AI world right now, as most other labs take 2-3 months for new models to be released.

Grok 4.6 is not a must use right now, but I would recommend checking it out if you have the chance to see how it works with what you do, with the expectation that we will be seeing more from them in the near future.

Quick Hits

Qwen 3.8 27B

The Qwen 3.6 series has been the go-to for people trying to run LLMs on local consumer hardware, and the latest Qwen 3.8 model perpetuates this.

The 27B model is a multimodal dense model, meaning that you will probably want a dedicated GPU to run it (sorry Mac users, although the mixture of experts variant is rumored to be coming out soon as well).

Qwen 3.6 27B benchmarks

I haven’t had much time to test it yet (why have labs started releasing models on Fridays?), but the initial vibe seems to be a fairly decent jumpp in agentic capabilities, but the limits for this have yet to be explored, for instance if it can replace DeepSeek V4 Flash as a capable subagent.

Of note, it has variable reasoning modes (from low to xhigh) similar to most other modern reasoning models. On the xhigh setting, it uses substantially more tokens than its predecessor, so be wary of that.

If you are running local models, you are most likely already using Qwen, and this seems to be a safe upgrade.

Finish

I hope you enjoyed the news this week. If you want to get the news every week, be sure to join our mailing list below.

DeepSeek V4 Pro self portrait — From xule on Twitter using Midjourney v8.2

Stay Updated

Subscribe to get the latest AI news in your inbox every week!

← BACK TO NEWS