Vector Lab
VECTOR LAB

EST. 2025

WEEKLY UPDATE2026
BY ANDREW MEAD

Jev

What the heck is Jev and open source models flop in WeirdML V3

Read in

Releases

Jev

A new startup called TypeSafe AI made waves this week with the release of their model called Jev. Despite getting a ton of hype as if it was an LLM, it actually isn’t.

Jev "benchmarks"

Jev benchmarked against LLMs on classification tasks. Note how the x-axis is logarithmic.

Instead it is what is known as a zero shot classification model. This means that you give it a piece of text and a set of categories, and it will give a probability for the text for each category.

These models tend to be very small, and because they just output probabilities, you only have to pay for the input tokens. This means that Jev only costs 4.2 cents per million input tokens, and you don’t pay for output tokens at all, since there are none.

This may seem revolutionary (as many people on Twitter think) but if you have been following AI since before the ChatGPT Cambrian explosion, you will know that this is not that revolutionary; we have been training these models (known as BERT models) since 2020!

We have a variety of models that already do this (like Modern Bert Zeroshot) and they are very easy to train (I was training them on my laptop back in 2021!). Unlike generative AI models, these more traditional ML models are much easier to train and deploy; they can easily run in your user’s browser, even on mobile.

Out of the box, Jev is a better version of what we have in open source right now and also outperforms the cheap LLMs that people have been using for classification tasks (like Gemini 3.5 Flash Lite and GPT 5.6 Luna), but its virality has spawned a number of clones already (many of which are open source) so I expect the gap to dwindle quickly.

It also should be noted that these are zero shot models, meaning they haven’t been trained for the specific classification task they are being used for. Like I said before, these models are very easy to finetune, requiring only a couple dozen samples to train, and can be done on most household hardware, no GPU required. These finetuned variants will vastly outperform the zeroshot models (and LLMs) while being effectively free for you to train and run.

That being said, if you have a classification pipeline and are using LLMs for it, I would replace them with Jev as it should be just as good if not better than most LLMs and vastly cheaper and quicker. Once you have done that, I would look into finetuning your own version to reduce your costs to 0.

Research

Weird ML V3

Speaking of finetuning models, a new finetuning benchmark has been released: Weird ML V3.

Prelim scores for Weird ML V3

X-axis is logarithmic. GPT-6 Astra gets the same score as Fable 5.1 with 1/10th the tokens.

The benchmark measures how well the model is able to do when dropped in to a machine learning environment with little direction outside its task. It is up to the model to look at the data, filter and process it, and then decide what architecture to use for the actual model, and then compose a pipeline to successfully train the model (they have only 2 minutes worth of compute to train with, so they must make sure their model is efficient).

The previous iteration was a good benchmark (Epoch AI recently benchmarked the benchmark, and found no issues with it, which is surprisingly rare) but was becoming saturated by the top models, so a new benchmark was in order.

The released 4 samples (while keeping the rest of the problems private) from the benchmark to get an idea of what they are looking for from the models. Tasks include shape detection and classification, ship detection, telescope data signal processing, and bonanza, which is all 17 tasks from Weird ML V2 that the model must solve in parallel, and the only score it gets is the average across all 17 questions, so it has to figure out itself which ones it is or isn’t doing will on.

The benchmark does a good job in differentiating between models, specifically open vs closed models. Open models tend to be trained around the major software engineering benchmarks and tasks, and tend to lack diverse coding capabilities, which most benchmarks miss. Weird ML does not have this issue, as there is a wide gap between most models, which gives a good measure generality.

For the preliminary scores, we see this, as Fable and GPT-6 Astra are far ahead, with GPT-6 in the lead due to its superior image understanding capabilities, since many of the problems benefit from strong multimodal capabilities. DeepSeek V4.1 Flash and Gemini 3.8 Flash, which seem close to the frontier on many benchmarks, are noticeably behind.

Weird ML V3 will be added to the list of benchmarks that I use for the news, which is nice to see, as many of the benchmarks I have relied on have become contaminated (like DeepSWE), but that is a discussion for another day.

Finish

I hope you enjoyed the news this week. If you want to get the news every week, be sure to join our mailing list below.

Doodles from Claude by Kevin Ngo on Twitter

Stay Updated

Subscribe to get the latest AI news in your inbox every week!

← BACK TO NEWS