- 1033comments
- 447comments
- 612comments
- 246comments
- 376comments
- 301comments
- 383comments
- 129comments
- 424comments
- 206comments
- 267comments
- 342comments
- 237comments
- 223comments
- 109comments
- 266comments
- 174comments
- 134comments
- 167comments
- 322comments
- 110comments
- 123comments
- 168comments
- 79comments
- 83comments
- 79comments
- 83comments
- 122comments
- 95comments
- 218comments
Hi HN. Jeff is a set of small, open-weight Qwen3.5 and Gemma fine-tunes for zero-shot classification, with respectable out-of-the-box performance, meant to be slotted right into code (or fine-tuned further as needed). You give them a situation and a list of options; they return a calibrated probability for each, in one forward pass, with no text generation. The 2B scores 83.1% on a five-benchmark panel (Jev's published figure: 83.0%); the 0.8B decides in about 28 ms on an M4 Max. Apache 2.0, with a Jev-compatible API (I'm not affiliated with TypeSafe).
When TypeSafe released Jev a couple of weeks ago and then AutoJev appeared, I wanted to see if I could replicate the experiment using only small language models on local hardware. So everything ran at home: one RTX PRO 6000 for training, two DGX Sparks running Qwen3.8-Flash-Next to write the synthetic data, a MacBook for testing, all monitored from my phone over Tailscale.
The caveat: the published Jev and AutoJev numbers are on a different sample of the same benchmarks, and Jeff's overall score comes from classification-style tasks (96% on Financial PhraseBank, 86-89% on RAGTruth, both above Jev). On multi-step reasoning it's behind: BBH 64-68% against Jev's 94%, and about 50% on JevBench's hard tier against 73%. That isn't surprising, and I don't think it matters: no 0.8B or 2B model reasons like a large one, and nobody should expect it to. These are extremely fast judgement-callers. In one of my apps I use the 0.8B for voice navigation; a quick fine-tune (about half an hour on one GPU) took it from 32% to 96% on held-out commands, at about 40 ms per decision.
The fun part: games, as a zero-shot test. Games aren't the ideal zero-shot test, but they're fun, and TypeSafe did it with Jev too. There was no game data in training. Each turn the code describes the situation and the moves in words, and the model picks one; the options say what each move leads to, never which one is right. Over 20 episodes each:
- Doom (ViZDoom): Jeff 0.8B 6.55 kills per episode, the same as a hand-coded bot and as Jev's published run. Jev's prompt spells out the aiming rule and takes about 212 ms per call over its API; Jeff gets "the nearest monster is a little to your left" and decides in about 29 ms on my Mac.
- Frogger: 10.3 crossings, level with the hand-coded bot (10.25), and 10x the untrained base model (1.0).
- Pac-Man: 57 of 98 pellets, about 60% of the bot's score and 2x the untrained model.
Videos of every run are linked in the README.
Lessons learned:
- System 1 models are here to stay. Being able to process unstructured data at software speed inside an app is extremely powerful, and being able to do it locally is fantastic.
- A small model is a classifier, not a planner. Models of 0.8B-2B don't reason like Qwen3.8-27B or Jev, and they don't need to: present the options the right way and you get 40+ decisions per second, depending on your hardware.
- Fine-tune it if needed. If zero-shot isn't good enough for your task, a short fine-tune on your own examples is.
- Wording matters enormously. Giving Frogger's final step the same words as every other forward option ("safe, and one row closer to the goal") took one episode from 15 crossings to 23. Before that, the frog just stayed on the last log.
- Bigger isn't better. The untrained 2B is already more risk-averse than the untrained 0.8B (in Doom it prefers turning away from the nearest monster), and training made it hesitate in Pac-Man. That's probably why the 0.8B beat the 2B.
- Benchmarks don't predict play. Untrained Gemma 4 E2B beats both untrained Qwens on the benchmarks (62.5%) and plays every game worst: right most of the time, but not reliably, and in a real-time loop the mistakes compound.
Funny that the 2B loses to the 0.8B. Question about the benchmarks: BBH and JudgeBench are more reasoning, where you fall behind, but for zero-shot classification there are more relevant ones like Banking77 or CLINC150. Was there no temptation to pick something closer to where System 1 models are actually used?
That's some local hardware.
Was going to make the same observation. Cool to run locally - but renting seems the saner choice?
Would love to see what this run would cost from something like Verda. Just out of curiosity. I'm not going to be installing any sparks at home anytime soon.
Great project! How does this compare to asking qwen to reply with 1 token in terms of speed?
This type of project looks extremely useful. There was a lot of buzz around Jev, but having models that run locally and can be fine-tuned is extremely helpful.
I compared it to Jev in my current use cases and it's very inaccurate. 70% vs 94% . for classification, it's unacceptable.
For _your_ classification it's unacceptable. The OP seems to have anticipated this and mentions you can fine tune it for your use case. Did you try that?
I don't think the point is to displace Jev, but to show it's possible to build an MVP on open weights without years of work and millions of dollars.
Why (presumably) an engineer would dismiss exploring a lightweight, custom alternative to locking into a fashionable PaaS, I'll never know.
Needing fine-tuning for the use-case completely changes the product category
Not sure the task at hand here. But if it doesn’t require any reasoning/thinking and it’s just a classification task, it’s worth a shot to look into training your own classifier
I’ve run some benchmarks. Using embeddings + logistic classifier, the architecture matches or beats Jev and Laya in all basic classification tasks (datasets tested: AG News, Emotion, MASSIVE Intent, Banking77) The type of task in which it does really well, especially against Laya, is classification with >50 classes
The classifiers also run in <1ms, so they can be very fast and precise at the same time
But this architecture has no “reasoning”, so it performs rather poorly on tasks that require it, like the ones from the XLNI dataset (Jev/Laya do a lot better on this one)
For the latter cases, you could use add a local lightweight LLM, something like a Gemma model. Or even some basic MLP, depending on the tasks/data
"Using embeddings + logistic classifier, the architecture matches or beats Jev and Laya in all basic classification"
Have I understood correctly that you trained only the logistic classifier, but didn't need to train the embedding model?
If so, I'm curious whether you compared that approach (A) with:
B) Jev only, with a single output.
C) Jev with multiple outputs fed into a logistic classifier.
Obviously C has cons (can't be self-hosted, needs some up-front work on deciding the shape of the output) but it might be somewhat more interpretable. (And I suppose it might have better performance?)
You are correct, I didn’t train the embeddings model
Here's a gist with code you can use to test the Banking77 dataset: https://gist.github.com/nicobrenner/056a5aaff5d0119c0032ecda...
The gist uses BAAI/bge-large-en-v1.5, which is 1.2GB approx. You can replace it for all-MiniLM-L6-v2 (91 MB @ fp32 or 45 MB quantized fp16) small enough for mobile/edge. With all-MiniLM-L6-v2 it still gets 93.0% on Banking77, only 1.3 points behind bge-large at 15x smaller
I haven’t compared different ways of sending requests to Jev
The data to train the classifiers comes from the datasets used to test them (not from Jev)
What do you use to determine that a particular task in a heterogeneous pile of tasks requires reasoning? The logistic classifier itself is too dumb to recognize the details of the problem that make it reasoning-sensitive (IIRC recognizing the “fiddliness” of a given problem requires a recognizer at least as complex as the problem itself.) And if you’re using the lightweight LLM for that, then you may as well skip the classifier and just use the LLM all the time, since that eval step is already going to be dominating your response time anyway.
My understanding of Jev is that it’s a replacement for the LLM you’d necessarily need to use to identify reasoning-sensitive workloads in a heterogeneous mix, where Jev will be cheaper than an actual LLM and so act as an actual optimization / de-bottlenecking change.
I’m in the process of piecing together the different task/dataset-specific classifiers
Depending on how much overfit, you can go from routing deterministically based on features/shape of the input data, all the way to training a routing model (which could be a classifier too). I’ll need to experiment to find the best approach
For completely unseen/unexpected, I’ve also experimented routing to a local LLM: request comes in, if there’s a marching classifier, send it there, otherwise send to LLM+training. As the system learns more tasks, the % of requests that go to the LLM go down over time
You can already so that with classification models such as ModernBERT, at 0.4B.
Jev's value is its zero shot performance without having to fine-tune.
I am sure I am missing something obvious here, but why is that valuable? Like, what kinds of projects are there where you need to classify stuff but are unable to make a bespoke model targeting the specific problem?
For my org, it meant we could trial classifiers across various internal systems with little to no engineering effort. In one case we ended up building our own classifier instead of Jev, but in others we kept Jev because it was zero-effort for a great impact.
My use case is simple classification for job ads. Things like, industry, work settings (remote, hybrid, onsite) and job type (full time, part time .. etc).
I did side by side comparison with Gemini 2.5 Flash Lite, Jev, Jeff
I tried the 0.8B model, completely useless in classification. Qwen Jeff-Qwen3.5-2B was better, but still missed job type.
I suppose with larger model, this could be useful, but would require more ram and will be slower.
Von 1.2 had a better Doom score :D
https://github.com/wfzyx/von
Oh wow, those Doom scores for Jeff are pretty terrible
The Von numbers have led me on a rabbit whole of getting a classifier to play Doom
I got it to average 22 kills (the max is 26) on that same scenario that Jeff and Von are testing on (it’s called Defend Center)
Now I’m having it play a more advanced scenario, and it’s doing about 45 kills (SOTA is ~59 kills)
It’s amazing what you can do with small classifiers if you can collect some data. These models I’m testing train on CPU in seconds (what takes the longest is running the game, doing test runs and collecting data), they are <1MB in size and do inference in <1ms on CPU
Edit: after looking at Jeff's numbers more in detail, the 6.5 kills number is not that bad, but it can definitely be better ;)
One day we’ll be using the same Kills/SOTA metric for models driving physical kill-bots.
Forgive my lack of understanding but how long before Jev type functionality is just built straight into all frontier models?
No need, as that functionality can run locally no problem.
the appeal would be if they can deliver it at a much higher performance and similar speed, which is plausible
BeRT and FLAN-T5 were used as classifiers 5-7 years ago, they were technically "frontier" for their time.
This is the exact comment I’ve been waiting for, what is the difference between classifiers and jev?
nothing in utility. we used various bert variants to satisfy our usecases and still in uses. free and they run locally.
Jev is a classifier. The big thing about it is that it has high accuracy on domains it wasn't fine-tuned for, like an LLM, but with speed and cost comparable to traditional classifiers.
BERT need to be fine tuned for your use case, Jev generalizes. It’s a pretty big difference!
FLAN-T5 generated text (Jev does not generate text), and BERT wasn't able to do tasks without fine-tuning.
Jev is basically a kind of FLAN-BERT, if you want, where it has built-in multi-task ability, but doesn't generate text. It only generates 255 floats all at once, making it much faster, and what those floats mean (if anything) depends on the prompt.
Eg, the following query is put in the encoder model:
{"question": "Rank these 5 things by increasing order of how big they are", "choices": ["truck", "cow", "mouse", "ant", "building"] }
The model returns [3., 2., 1., 0., 4.], and 249 other meaningless floats that are hidden from you by the UI.
The UI stitches the first 5 floats with the choices and returns something like:
{"rank": ["ant", "mouse", "cow", "truck", "tower"]}
So it generates logits in a 255 token output space? ;)
logits assumed some form of softmax or logistic, which may not be the case
How does it know that you’re asking for “rank” instead of something else if it’s not generating text?
If I had to guess, it's already built and is just waiting on Product's/Marketing's desk. How do you position this without looking like your roadmap is being determined by newcomers? Probably don't want to adopt the same verbiage+acronyms - but also can't be seen to be just sherlocking features.
I think Apple has demonstrated that shipping second has essentially no negative impact if your product is seen as higher quality.
I think Nvidia has demonstrated that shipping first is a multi-trillion dollar opportunity if you don't shy away from a challenge.
I think it depends on the model: Nvidia ship products with APIs, apple/jev/etc. ship end user products. The former is much more subject to lock-in, increasing the value of early market adoption because there's a 3P Nvidia ecosystem sprung up in response. Apple/jev/etc. do have APIs & corresponding 3P ecosystems but those are usually a smaller component of market capture than direct product end users, so the space ends up more competitive.
Chip manufacturing intrinsically comes with one hell of a moat. There's not much parallel in software.
That’s kind of funny because Nvidia’s biggest moat is arguably CUDA, the software ecosystem around their chips
CUDA is a lock-in moat, the infrastructure needed for chip manufacturing is a barrier-to-entry moat. Two different things.
There are many microarchitecture patents used in NVIDIA chips. I'm sure they have a team that rips apart AMD chips looking for infringement. The way CUDA works is tied to many GPU architecture decisions and it would be hard to decouple them efficiently. Obviously a huge effort was made to get PyTorch decoupled from CUDA.
If cuda is an API, and llms make apis effortless, then how big of a moat is cuda?
Historically there was a big patent moat in (Graphics) GPU design. This continues with CUDA, but obviously Intel and AMD could find ways to support eg PyTorch. What we don't know is how much effort that cost them, or why they decided they couldn't make a CUDA API compatible competitor.
NVIDIA literally hired all of the 3DFX team and patents after they already proved the graphics accelleration hardware market was massive with the Voodoo card line
In fact NVIDIA wasn't even a close competitor to 3DFX in the graphics card game at that point
It panned out great. 3DFX filed for bankruptcy less than 18 months later, and Nvidia could pivot from designing raster chips to considering CUDA's architecture.
It's not like 3DFX was the first GPU vendor. Nvidia saw the opportunity to be the first true GPGPU vendor, and they beat their competitors.
It's more due to ineptitude of competition. If AMD shipped second, but better product it would be another thing but it is still a bit of a mess of an ecosystem on AMD side
I don't think there would be any utility for that. Anything jev can do, a frontier model can also do. Just not as quickly or as cheaply.
I think these products (Jev and the inevitable offerings from Anthropic, OpenAI, etc) want to become more than end-user output machines. They'd benefit from being in the hotpath of other services. Not backgrounded generation but in-band, request-time work.
1M x $0.50 == 1B x $0.0005
To expand a bit for my current use cases. Inline routing of work to heavy task specific models, and prompt/context generation (user is asking something, what and how much should we prompt the expensive LLM with). Latency or time to first token does matter for some applications.
Probably at the frontier stage - you will only see it where Jev is better regardless of cost.
For everyone else who is conscious of cost, you're already seeing this being built into harnesses.
Almost certainly, you'll see versions of this from all the Chinese labs as fast as humanly possible.
If I had to guess, Cursor/Grok or Google/Antigravity will be the first major players to natively support something like this to drive down cost, as they're primarily the budget conscious choices.
I would be astounded if Anthropic leads the way on a cost reduction.
But for those of us that prefer open source and self hosting, JEV alternative LAYA will beat anything the frontier models package up.
Negative three years, give or take. Although recent Anthropic and OpenAI models no longer expose the capability. But for any open model you just tell it to respond with a single token "Y/N" and take the logit difference. If you want multiple distinct questions answered you just ask them independently and put the shared context first so it gets cached.
OpenAI and Anthropic don't want to give out logprobs these days but could trivially add a dedicated classification API to their existing models if there was enough demand.
I think the main differentiator offered by Jev is not the ability to answer questions, most models can be coerced into that function if they don’t already have a dedicated pipeline for it, rather it’s the extreme speed of the evaluation, and very low cost that opens new possibilities.
The evaluation is fast because it's all prefill computation with only a single token of inference. Ditto cost, you're paying 100% input costs and nearly zero output. There really isn't any architectural magic to Jev, it's just a straightforward application of normal LLM tech with some good marketing.
I mean, Jev is also probably cheaper because it's a rather small model (or at least, I suspect it is based on the overall level of intelligence it demonstrates) so that helps make it cheap too.
Yes this is exactly my assumption as well. I think ppl forgot before agents it was expected slo to have a ttft in range of a few hundard ms, which is what jev is also achieving.
Thats also why it feels weird they say they don't "charge for output tokens" since its literally generating a single (or at most very few tokens).
Can we get a price comparisson ?
Edit: Running them for the masses.
I assume the cost is whatever you run the model on. Jev is already dirty cheap, $0.42 per million input tokens and output is free and the input cost is covering tokenization + API utilization
Well, I'm no expert, but this runs on Qwen3.5 and Gemma 4, stripped down, but they're pretty pricey.
$0.042 not $0.42
Typesafe has been quite about the underlying technology behind Jev. Given the speed and cost my hypothesis is that it doesn’t input tokens the way that LLMs do, ie iterating over every word and drawing the connections between each. That is an o(n^2) problem which is why LLMs are so expensive as they scale.
Most likely: it does still have attention layers (the O(n^2) part), but it’s not autoregressive (which makes it O(n^3) because you have to run the whole model again for each predicted token)
Then again, it's only a very small number and fixed set of tokens for the output.
I feel like the most likely is that it largely works like how all the recent copycats work: take an existing llm, modify the decoder, do RL training. I think the main reason that jev works better is that they spent more time on that post training step.
My guess as well
The main difficulty with fast (low-latency) inference, is not actually computation but loading parameters from memory. The problem with generating 1 token at a time isn't that that its expensive computationally (it is, but so is training), but that you need to stream your entire model from memory for every single token. So the strategy here is to process big batches of user requests, letting you share the memory loads across users. This means individual answers aren't that fast, but you can do lots at once.
My hypothesis for Jev is that they simply generate many answers independently in parallel from your prompt, and then discard the duplicates (or train to avoid duplication). In that way the entire batch is 1 user's prompt.
What proportion of commercial LLM use is classification? I'm just wondering what happens to business AI spending/data centre usage when they realise they don't need full LLMs.
My name jeff
Anything like this in the VLM side? Classification on images...
Pretty straightforward: https://news.ycombinator.com/item?id=49853175
Isn't jev just a less nuanced classifier? What am I missing?
The zero-shot learning is what makes these different from a traditional classifier.
Zero-shot just means you give it zero examples. Jev lets you add examples, so it's one-shot/few-shot depending on how many you provide. After playing with it, it seems like once you wonder off their guide examples domains - you have to provide examples to get anything useful from it.
IMO you going to get better results by doing a small tune of a tiny model. Making training dataset for it with LLMs is easy, serving it is going to be cheaper.
I've always been more afraid of these types of models than LLMs. These are what enable mass surveillance at scale and autonomous real time combat drones. Now they are spreading and being optimized. Gg.