Vector Lab
VECTOR LAB

EST. 2025

WEEKLY UPDATE2026
BY ANDREW MEAD

Kimi K3

Kimi joins the frontier and is frontend really solved?

Read in

Releases

Kimi K3

Moonshot AI has released the latest generation of their Kimi series of models, Kimi K3. The Kimi models were already some of the largest open source models that we have seen, being 1 trillion parameters, and now they have 3x’d this to 3 trillion parameters for their 3rd gen flagship, double the estimated size of GPT 5.5 and Opus 4.8.

Kimi K3 benchmarks

The model is by far the best we have seen from China, and in fact the best from any lab that isn’t Anthropic or OpenAI. I would place the model right behind the frontier of Fable and GPT 5.6, but is close behind. It does not have the same high quality judgment ability of Fable, and lacks the extreme token and tool use efficiency of GPT 5.6.

It has a less refined feel and user experience compared to the top two, but in terms of intelligence it seems to be close, making these issues very easy to overcome in future iterations of the model, and they can be ironed out in post training.

Similar to other Chinese models, Kimi K3 excels in frontend design, topping GPT 5.6 and Fable in the Frontend Code Arena. On the very reliable DeepSWE benchmark we see that it is behind only Fable and GPT 5.6 Sol, further reinforcing the vibes that I and many others have been seeing in the real world that this model is one of the best in the world.

The model is also multimodal, but similar to the Claude models, it does not seem that the Kimi team has focused too much on these capabilities compared to OpenAI.

One of the big departures from previous Chinese releases is the price. They have decided to price it at $15 per million output tokens, which is the first time any of the frontier Chinese labs have priced a model above $5 per million output tokens. This shows that they are confident in the model’s capabilities and that they can start taking a much larger profit than they have previously.

Model$ per million (input)$ per million (output)Tokens per second
Kimi K3$3$1527
GPT 5.6 Sol$5$3045
Claude Fable 5$10$5040
Claude Opus 4.8$5$2560
GLM 5.2$1.40$4.4047

Right now, the model’s high price and high token usage makes it both slower and more expensive to use than GPT 5.6 Sol. We expect to see these prices come down however, as the model will be open sourced next week, and all of the 3rd party inference providers will start supporting the model and trying to undercut each other on price.

Also due to its open source nature, you will not have to deal with the same guardrails that you do with GPT and Claude. You will get to use the model however you want, and won’t have to worry about getting blocked for asking about cyber security or biology related tasks.

In just one month we have gone from GLM 5.2 which gave us very usable but not necessarily frontier level capabilities to Kimi K3, which is a roughly frontier LLM. It is clear that OpenAI and Anthropic will not run away from everyone else and control the entire AI space. Now it is not even clear if they will be the top labs at the end of the year. OpenAI and Anthropic had a head start on everyone else, but now that other labs are catching up, the AI race has truly begun.

Quick Hits

React Bench

Many of us who use LLMs for coding have seen frontend tasks as a mostly solved science, as LLMs are capable of making very nice looking frontends for our websites with minimal effort.

React Bench challenges this notion. It measures bugs and performance issues with LLM-made code, and it finds that even frontier models still struggle. They test both greenfield tasks and also bug fixing tasks across 51 different scenarios.

Users notice bugs far more than features and are very sensitive to the snappiness of your frontend. The frontend is also usually the hardest for your LLM to test, since it needs to interact with the browser using a tool like playwright and take screenshots to understand what is happening, and LLM’s vision capabilities are far behind their text capabilities.

Finish

I hope you enjoyed the news this week. If you want to get the news every week, be sure to join our mailing list below.

Floating Companion — from ClankrMedia on Twitter

Stay Updated

Subscribe to get the latest AI news in your inbox every week!

← BACK TO NEWS