Hacker News

Top stories

Live mirror
30 storiesupdated just nowView source snapshot
  1. Jeff – Jev-compatible 0.8B decision models, trained at home, ~30 ms (github.com/firelex)
    34comments
  2. Pirating the Pirates (mubi.com)
    181comments
  3. MicroLLM Lab – Try 7 tiny LLM's in the browser (stateofutopia.com)
    42comments
  4. 12,000-year-old Göbeklitepe burials explain scattered bones (archaeologymag.com)
    7comments
  5. World Labs Is Joining AMD (worldlabs.ai)
    45comments
  6. Scientists solve 1840s space weather mystery (arstechnica.com)
    20comments
  7. Who should be held accountable when an AI Agent (accidentally) acts maliciously? (greenpants.net)
    8comments
  8. Hijacking the PS5's RTMP stream (yashgarg.dev)
    54comments
  9. Joseph Szabo’s pictures of American adolescents (newyorker.com)
    32comments
  10. It's Time to Investigate the AI Labs (calnewport.com)
    46comments
  11. Kids turned low-traffic NPR Spotify comments into a secret group chat (thisamericanlife.org)
    141comments
  12. Parley: Federated, decentralised chat that speaks plain IRC (mills.io)
    145comments
  13. Sonnet 5.5 (anthropic.com)
    339comments
  14. Flock Wants the Most Detailed Map of Its Surveillance Cameras Taken Offline (theintercept.com)
    31comments
  15. Best of British Design (best-of-british-design.vercel.app)
    21comments
  16. Pacing the Frontier is not the actual goal for AI labs (lesswrong.com)
    51comments
  17. Cf: The Agentic CLI for the Cloudflare API (cloudflare.com)
    40comments
  18. What reversing, modernising old games tells us about the economic impact of AI (isfine.org)
    8comments
  19. First Steps of the PLC Organization – Independent Public Ledger of Credentials (plcred.org)
    18comments
  20. 3D necroprinting: Leveraging biotic material as the nozzle for 3D printing (science.org)
    —discuss
  21. Nvidia wants to put a watchdog chip next to every AI agent (cnbc.com)
    123comments
  22. Launch HN: Vespper (YC F24) – SOTA Docx MCP (vespper.com)
    8comments
  23. Show HN: Destroy Any Website with Stickman (spritefusion.com)
    25comments
  24. Who wrote Elizabeth I's most scathing letters? (smithsonianmag.com)
    18comments
  25. Does Reddit have an astroturfing problem? What the data suggests (petervijeh.com)
    80comments
  26. What heraldry and Japanese mon can teach about visual-identity generators (benovermyer.com)
    19comments
  27. Updated Google Maps shows destruction of the city of Rafah (twitter.com/aliabunimah)
    34comments
  28. MongoDB CEO resigns to join Meta (reuters.com)
    252comments
  29. GrapheneOS – When an app is slow (wirelessmoves.com)
    38comments
  30. I made a visual workspace for AI Automations (biom.dev)
    22comments

Jeff – Jev-compatible 0.8B decision models, trained at home, ~30 ms

133 pointsby 2h agogithub.com
29 comments
2h agoHN ↗

Hi HN. Jeff is a set of small, open-weight Qwen3.5 and Gemma fine-tunes for zero-shot classification, with respectable out-of-the-box performance, meant to be slotted right into code (or fine-tuned further as needed). You give them a situation and a list of options; they return a calibrated probability for each, in one forward pass, with no text generation. The 2B scores 83.1% on a five-benchmark panel (Jev's published figure: 83.0%); the 0.8B decides in about 28 ms on an M4 Max. Apache 2.0, with a Jev-compatible API (I'm not affiliated with TypeSafe).

When TypeSafe released Jev a couple of weeks ago and then AutoJev appeared, I wanted to see if I could replicate the experiment using only small language models on local hardware. So everything ran at home: one RTX PRO 6000 for training, two DGX Sparks running Qwen3.8-Flash-Next to write the synthetic data, a MacBook for testing, all monitored from my phone over Tailscale.

The caveat: the published Jev and AutoJev numbers are on a different sample of the same benchmarks, and Jeff's overall score comes from classification-style tasks (96% on Financial PhraseBank, 86-89% on RAGTruth, both above Jev). On multi-step reasoning it's behind: BBH 64-68% against Jev's 94%, and about 50% on JevBench's hard tier against 73%. That isn't surprising, and I don't think it matters: no 0.8B or 2B model reasons like a large one, and nobody should expect it to. These are extremely fast judgement-callers. In one of my apps I use the 0.8B for voice navigation; a quick fine-tune (about half an hour on one GPU) took it from 32% to 96% on held-out commands, at about 40 ms per decision.

The fun part: games, as a zero-shot test. Games aren't the ideal zero-shot test, but they're fun, and TypeSafe did it with Jev too. There was no game data in training. Each turn the code describes the situation and the moves in words, and the model picks one; the options say what each move leads to, never which one is right. Over 20 episodes each:

- Doom (ViZDoom): Jeff 0.8B 6.55 kills per episode, the same as a hand-coded bot and as Jev's published run. Jev's prompt spells out the aiming rule and takes about 212 ms per call over its API; Jeff gets "the nearest monster is a little to your left" and decides in about 29 ms on my Mac.

- Frogger: 10.3 crossings, level with the hand-coded bot (10.25), and 10x the untrained base model (1.0).

- Pac-Man: 57 of 98 pellets, about 60% of the bot's score and 2x the untrained model.

Videos of every run are linked in the README.

Lessons learned:

- System 1 models are here to stay. Being able to process unstructured data at software speed inside an app is extremely powerful, and being able to do it locally is fantastic.

- A small model is a classifier, not a planner. Models of 0.8B-2B don't reason like Qwen3.8-27B or Jev, and they don't need to: present the options the right way and you get 40+ decisions per second, depending on your hardware.

- Fine-tune it if needed. If zero-shot isn't good enough for your task, a short fine-tune on your own examples is.

- Wording matters enormously. Giving Frogger's final step the same words as every other forward option ("safe, and one row closer to the goal") took one episode from 15 crossings to 23. Before that, the frog just stayed on the last log.

- Bigger isn't better. The untrained 2B is already more risk-averse than the untrained 0.8B (in Doom it prefers turning away from the nearest monster), and training made it hesitate in Pac-Man. That's probably why the 0.8B beat the 2B.

- Benchmarks don't predict play. Untrained Gemma 4 E2B beats both untrained Qwens on the benchmarks (62.5%) and plays every game worst: right most of the time, but not reliably, and in a real-time loop the mistakes compound.

1h agoHN ↗

Funny that the 2B loses to the 0.8B. Question about the benchmarks: BBH and JudgeBench are more reasoning, where you fall behind, but for zero-shot classification there are more relevant ones like Banking77 or CLINC150. Was there no temptation to pick something closer to where System 1 models are actually used?

31m agoHN ↗

Was going to make the same observation. Cool to run locally - but renting seems the saner choice?

Would love to see what this run would cost from something like Verda. Just out of curiosity. I'm not going to be installing any sparks at home anytime soon.

1h agoHN ↗

This type of project looks extremely useful. There was a lot of buzz around Jev, but having models that run locally and can be fine-tuned is extremely helpful.

1h agoHN ↗

I compared it to Jev in my current use cases and it's very inaccurate. 70% vs 94% . for classification, it's unacceptable.

47m agoHN ↗

For _your_ classification it's unacceptable. The OP seems to have anticipated this and mentions you can fine tune it for your use case. Did you try that?

I don't think the point is to displace Jev, but to show it's possible to build an MVP on open weights without years of work and millions of dollars.

Why (presumably) an engineer would dismiss exploring a lightweight, custom alternative to locking into a fashionable PaaS, I'll never know.

40m agoHN ↗

Needing fine-tuning for the use-case completely changes the product category

35m agoHN ↗

Not sure the task at hand here. But if it doesn’t require any reasoning/thinking and it’s just a classification task, it’s worth a shot to look into training your own classifier

I’ve run some benchmarks. Using embeddings + logistic classifier, the architecture matches or beats Jev and Laya in all basic classification tasks (datasets tested: AG News, Emotion, MASSIVE Intent, Banking77) The type of task in which it does really well, especially against Laya, is classification with >50 classes

The classifiers also run in <1ms, so they can be very fast and precise at the same time

But this architecture has no “reasoning”, so it performs rather poorly on tasks that require it, like the ones from the XLNI dataset (Jev/Laya do a lot better on this one)

For the latter cases, you could use add a local lightweight LLM, something like a Gemma model. Or even some basic MLP, depending on the tasks/data

43m agoHN ↗

Oh wow, those Doom scores for Jeff are pretty terrible

The Von numbers have led me on a rabbit whole of getting a classifier to play Doom

I got it to average 22 kills (the max is 26) on that same scenario that Jeff and Von are testing on (it’s called Defend Center)

Now I’m having it play a more advanced scenario, and it’s doing about 45 kills (SOTA is ~59 kills)

It’s amazing what you can do with small classifiers if you can collect some data. These models I’m testing train on CPU in seconds (what takes the longest is running the game, doing test runs and collecting data), they are <1MB in size and do inference in <1ms on CPU

53m agoHN ↗

Forgive my lack of understanding but how long before Jev type functionality is just built straight into all frontier models?

50m agoHN ↗

No need, as that functionality can run locally no problem.

46m agoHN ↗

BeRT and FLAN-T5 were used as classifiers 5-7 years ago, they were technically "frontier" for their time.

35m agoHN ↗

This is the exact comment I’ve been waiting for, what is the difference between classifiers and jev?

32m agoHN ↗

nothing in utility. we used various bert variants to satisfy our usecases and still in uses. free and they run locally.

20m agoHN ↗

Jev is a classifier. The big thing about it is that it has high accuracy on domains it wasn't fine-tuned for, like an LLM, but with speed and cost comparable to traditional classifiers.

15m agoHN ↗

BERT need to be fine tuned for your use case, Jev generalizes. It’s a pretty big difference!

44m agoHN ↗

If I had to guess, it's already built and is just waiting on Product's/Marketing's desk. How do you position this without looking like your roadmap is being determined by newcomers? Probably don't want to adopt the same verbiage+acronyms - but also can't be seen to be just sherlocking features.

36m agoHN ↗

I think Apple has demonstrated that shipping second has essentially no negative impact if your product is seen as higher quality.

30m agoHN ↗

I think Nvidia has demonstrated that shipping first is a multi-trillion dollar opportunity if you don't shy away from a challenge.

22m agoHN ↗

I think it depends on the model: Nvidia ship products with APIs, apple/jev/etc. ship end user products. The former is much more subject to lock-in, increasing the value of early market adoption because there's a 3P Nvidia ecosystem sprung up in response. Apple/jev/etc. do have APIs & corresponding 3P ecosystems but those are usually a smaller component of market capture than direct product end users, so the space ends up more competitive.

19m agoHN ↗

Chip manufacturing intrinsically comes with one hell of a moat. There's not much parallel in software.

26m agoHN ↗

The confidence scores need to be good if we're gonna forego fast and cheap.

38m agoHN ↗

Can we get a price comparisson ?

Edit: Running them for the masses.

36m agoHN ↗

Typesafe has been quite about the underlying technology behind Jev. Given the speed and cost my hypothesis is that it doesn’t input tokens the way that LLMs do, ie iterating over every word and drawing the connections between each. That is an o(n^2) problem which is why LLMs are so expensive as they scale.

31m agoHN ↗

Most likely: it does still have attention layers (the O(n^2) part), but it’s not autoregressive (which makes it O(n^3) because you have to run the whole model again for each predicted token)

10m agoHN ↗

Then again, it's only a very small number and fixed set of tokens for the output.

15m agoHN ↗

What proportion of commercial LLM use is classification? I'm just wondering what happens to business AI spending/data centre usage when they realise they don't need full LLMs.