Hacker News

New stories

Live mirror
30 storiesupdated just nowView source snapshot
  1. Microsoft Systems Journal Interviews Gordon Letwin (1987)(computeradsfromthepast.substack.com ↗)
    discuss
  2. Been thinking about our platforms and information diet we consume lately
    discuss
  3. How to export ChatGPT to Google Docs without losing formatting [video](youtube.com ↗)
    discuss
  4. AI is reaching new highs with AROM Labs
    discuss
  5. On Contagion (2019) – Boris Cherny(borischerny.com ↗)
    1comments
  6. ZuckOff Know when a camera is in the room(zuckoff.app ↗)
    discuss
  7. Forward Deployed Engineers from the Trenches(javisantana.com ↗)
    discuss
  8. Are We Free?(unixdad.com ↗)
    discuss
  9. Show HN:Linux2Win Auto-detect Linux disks, project Ext2/3/4 as Windows folders(github.com/aps-ppr-2016 ↗)
    discuss
  10. Is your translation production-ready: unit-testing your text(languageops.com ↗)
    discuss
  11. Show HN: Open Source AI Employees(github.com/markfulton ↗)
    discuss
  12. ZuckOff Is a Free App That Sees Meta Glasses Before They See You(wired.me ↗)
    discuss
  13. Pixel Area, a new directory of personal and independent websites(pxlarea.com ↗)
    1comments
  14. How the EU's age-verification app for children would work(reuters.com ↗)
    discuss
  15. Skia Compositor for WPE WebKit and WebKitGTK(igalia.com ↗)
    discuss
  16. A thing we may be able to learn from AI(wilsoniumite.com ↗)
    discuss
  17. Zep.js: tiny event pipeline with auto-cleanup, debounce and AbortSignal(github.com/marsbos ↗)
    1comments
  18. Agent Needs an Unknown State(drjoshcsimmons.com ↗)
    discuss
  19. Security: PolinRider malware detected in two open PRs (#7716, #10321)(github.com/shadcn-ui ↗)
    discuss
  20. Digital euro makes debut in wholesale financial markets(ft.com ↗)
    discuss
  21. Weather2 – new SQL datasets and multilingual dashboard
    discuss
  22. How to Smash the Memory Wall Plaguing High Performance Systems(nextplatform.com ↗)
    discuss
  23. Thomson Reuters: Thomson-1.0-Small(huggingface.co ↗)
    discuss
  24. Show HN: Deskies – ambient presence for friends across companies(deskies.vercel.app ↗)
    discuss
  25. AI Poster Prompts Improved(john.hartnup.uk ↗)
    3comments
  26. Frontier Overhangs(stratechery.com ↗)
    discuss
  27. New air traffic control failure delays flights at UK airports(reuters.com ↗)
    discuss
  28. You have more influence over the future than you think(sharif.io ↗)
    discuss
  29. Quantum Gravity, UAPs, Wormholes(twitter.com/jacksarfatti ↗)
    discuss
  30. Linker I/O tricks and their downsides(maskray.me ↗)
    discuss

Kev: Tiny Jev-like family of decision models built on top of Qwen3.5

103 pointsby 3h agogithub.com
50 comments
3h agoHN ↗

Been hoping for something in this space. Jev-like decision models on Qwen3.5 could really simplify some of our internal routing logic.

3h agoHN ↗

Quite impressed by the energy people are putting into making OSS Jev-like models.

I understand the hype but I wonder: what are the use cases for this kind of model? Could it be used in the context of coding agents, or is it more relevant in totally different situations?

3h agoHN ↗

Yeah same. Got access to their API and then realised I don’t really have an immediate use case

2h agoHN ↗

You should call Jev-like models when you give it a JSON-like structure to produce, it is useful when you need _some_ intelligence in your code.

Edit: I want to add that you can see Jev like a smart if-statement.

2h agoHN ↗

To develop a smart ai system for my 2d roguelike platformer? game has way too many moving system for classic state-machine ai + i cant spare the time to develop it. its low latency entices me.

2h agoHN ↗

Consider every situation where you "force" an LLM to output only a choice / category, or a set of them. If you have workflows like that, you're now being promised significant cost- and latency reduction.

For coding agents it'd only be useful in a subset of situations. E.g. you could imagine using one to classify bash tool calls into safe and unsafe for example.

2h agoHN ↗

what are the use cases for this kind of model? Could it be used in the context of coding agents

Yeah, it could. The most obvious usage would be to have local fast cheap "feedback" / "control" over a slower more expensive agent (i.e. cc / codex / opencode). Things like "goals" could now be split from a long prompt into "actions" and "verifiers". Where for each action you also produce a verifier. Then after each action you run the verifier w/ this kind of "universal classifier" and decide if the step was done correctly, if it needs follow-up and so on.

Example: implement auth in this repo -> llm_plan() -> for item in plan generate_verifier() -> for item in plan implement() ; verify() ; accept() / followup().

Verifiers could be something like this. take a plan item as input, generate classification questions that might verify the task "is this following project conventions?" | "is this touching files from other tasks?", etc.

You can do that with LLMs, but some things might become cheaper / faster. And you can pretty much use it to check against an ever growing list of conventions. Yours or project specific.

14m agoHN ↗

I don’t think this is a very good use case.

You can already do this much better with a normal LLM.

The problem with doing it with a model like Jev is that it’s an extremely dumb model.

2h agoHN ↗

Bit of a Jev explosion going on. Is it because it's taking us back to a simpler time we understand better? Classification models have been around for a while.

2h agoHN ↗

It’s because it’s practically useful and enabled things that were impractical previously.

1h agoHN ↗

and enabled things that were impractical previously

I think that there are not _that_ many use-cases that have been opened up by this that tool-calling on other models didn't solve already. Really depends what benchmark you're looking at. This one against BANKING77[0] has many issues, but suggests it's really not far off DeepSeek 4.1 Flash. This one against BoolQ[1] shows marginal improvement over Qwen3.6. This one against MMLU-Pro[2] (same author as the previous) shows significant improvements over two Qwen models.

So there's definitely _some_ alpha there, but I don't think it's the sea-change that the hype would suggest; that is to say, yes, some things that weren't practical before are now, but many things were already very practical with the existing tools.

0: https://sanand0.github.io/llmevals/jev/

1: https://github.com/ekzhang/openjev-sglang/blob/a3554ed9e9c26...

2: https://github.com/ekzhang/openjev-sglang/blob/a3554ed9e9c26...

2h agoHN ↗

It's appealing not having to fine-tune separate model for each use case

So you have more flexibility to get on with building, evolve your business logic etc

1h agoHN ↗

It reminds me a bit of what Ansible got right: user communication. The underlying tech may have existed for a long time, but the genius is presenting it to a regular developer in a way that reads "yes, even you can understand ML, just using a little JSON". The contribution of that should not be understated, as has been clearly evident recently.

33m agoHN ↗

Yeah except it doesn't really work. It constantly breaks underneath you. The whole system has to be managed, e.g NixOS, or else it's a house of cards.

7m agoHN ↗

I'm also not personally a fan of Ansible, but to claim it doesn't really work is quite breathtaking given the size of the installed base.

1h agoHN ↗

The way I see this (I havent played around with Jev or layla the OSS version) is that classifiers have always existed and a recognised tool in the ML world. But, the norm is that one needs to not only know what to classify as, but determine what weights to use to classify the input.

Jev came in, and added that magic of "you dont need to train your classifier or determine the weights" if you dont want to, and just get the classified answer out. I think that's what is making people see this with a glitter in their eyes.

39m agoHN ↗

I would be curious to see comparisons of jev and similar things with problem specific classifiers. I think layla suggested making problem specific versions anyway? There is a lot of demand for magic don't do any work solutions, which is kind of weird in an era where agents can really help you build a customised solution effectively.

17m agoHN ↗

Agreed.

Just to be helpful if anyone is searching for layla, it's laya.

1h agoHN ↗

Jev is creating a sort of identity crisis for me, because the number of absolutely clueless folks parroting the classifier thing is the first time I've seen this sort of mass psychosis in CS upfront.

Like even 5 minutes of tinkering captures why this isn't anymore like BERT or any past classification model than ChatGPT is like those old Markov Chain generators, yet folks cannot shut up about how this is nothing new.

Absolutely scary and makes me wonder how much of the field is just people super confidently discrediting otherwise promising/interesting directions for development for a cheap dunk!

1h agoHN ↗

Hey I am clueless, how do I learn more?

Why is Jev fundamentally better than classification models like BERT or traditional ML?

Happy to read a written response or if you suggest a prompt to put into my LLM to get it to research and explain the relevant details.

35m agoHN ↗

You already wrote the prompt, no? What I'd do, if I were you, is run the question through a LLM and then come back with targeted questions that it didn't answer.

I did the first part yesterday, jumped down the rabbit hole, and have 3 product ideas in my head now.

"Why is Jev fundamentally better than classification models like BERT or traditional ML?"

1h agoHN ↗

For a while is the keyword. It’s just vibe coders have just discovered the classifiers

2h agoHN ↗

Why does nobody ever ship these as a docker image?

2h agoHN ↗

I guess you have AI to write your docker files and push your images now.

2h agoHN ↗

I think a great use case for these will be when they have large context windows and are able to enforce styling rules for frontend development, and component creation rules for react. You can then ditch the styles guides and styling skills and create a decision tree for enforcing styling, so that you can't run into drift issues or duplication issues. That's where I'm wasting most of my time right now, constantly correcting all of the UX/UI issues that are created for every single feature.

1h agoHN ↗

Back in the good old days we would prevent these ux/ui issues by rigorously enforcing the use of our own stylesheets and classes. Later that grew to only using the company ux components. This was very successful in keeping everything neat and tidy. The only drawback was creating and curating new elements and getting consensus. But otherwise it works wonders.

Try constructing reusable components out of what you are doing instead of building everything up from basic building blocks. This also allows more concrete testing of individual parts and then if you want to change the look you can change it in one place and have it apply everywhere.

Agentic development doesn’t mean “throw all what we learned out of the window”, the same practices that helped speed up and improve quality of work of humans also helps agents. In fact, the multiplier is even bigger. You will notice it in development speed and reduced cost due to avoiding churn.

1h agoHN ↗

Hi, I'm currently planning on launching a product like this "Grammarly for Design" in a few weeks. Would you be interested in being part of the alpha group ?

2h agoHN ↗

Because these decision models do not have tool calling, the knowledge cutoff might become a problem. We'll either have to keep training continuously if we run locally or switch to the newer version every month or so when using a closed one like Jev

2h agoHN ↗

even with knowledge cutoff set a second from now, you still want to provide as much info as you can if you’re using such tools for delegating decisions

2h agoHN ↗

I wonder how these would do filtering my spam. I have been using 27B-class models for a while now, and they are nearly perfect at determining what is spam and what isn't. The only disadvantage is computational cost.

2h agoHN ↗

Take a look at Thomson 1.0-small, which is a variant of qwen 3.6 35b post trained by Thomson Reuters for text analysis. It classifies text content very well.

2h agoHN ↗

On Gemma 4 12B, I am getting 220 ms per move or QS. I used it to play the Snake game locally:

prompt_eval=244 ms wall=245 ms schema_cache=hit generated=0

Move limit reached after 200 moves: score=16, length=19.

So, if a 12B dense model can offer this latency on a local old PC, then definitely you can scale it up with more powerful machines and get even lower latency.

2h agoHN ↗

Can someone tell me what is the difference between Jev and a normal neural network that does classification ?

My understanding is: it takes text input and it does one shot classification (no training data)

2h agoHN ↗

Yes, this is essentially it.

As a corollary, the output classes can be any set, rather than needing to be set before training.

1h agoHN ↗

Can someone do a ELI5A of how they achieve classification over any user defined list of items? Normal neural networks do a softmax over a known output set to get probabilities

1h agoHN ↗

I can think of two possible approaches 1. Jev limits to 255 distinct options. So they can preprocess your set of options and “tell” the LLM via input tokens 1 = red, 2 = blue, etc then jev need only output softmax over 255 states while benefiting from pretrain of other LLMs 2. You allow the forward pass to output over the total token state but mask over the logits to limit to the user options. Less plausible? bc tricky when input is multi token which they clearly support.

My guess would be option 1. Didn’t read the kev repo here which would also explain

23m agoHN ↗

You can achieve open-vocabulary classification by making the final weights in the softmax come from a category encoder instead of being fixed learned weights. So instead of

softmax(encode(input)*learned_weights)

You have

softmax(encode(input)*encode(categories))

I'm not sure if Jev does it this way, but it's how you get open-vocabulary zero-shot image classification with models like CLIP [1].

[1] https://openai.com/index/clip/

2h agoHN ↗

Interesting approach with Qwen3.5 for decision models. Curious how "tiny" they've made them while keeping LLM reliability for critical paths.

2h agoHN ↗

Interesting to see a Jev-like approach applied to Qwen3.5. Always appreciated Jev's simplicity for quick decisions.

1h agoHN ↗

The bright side of Jev being so popular could be that many companies and individuals realize that their applications might work well with a System One model, and they decide to run an open-source (or fine-tuned) version on their own

1h agoHN ↗

Oh Jared is cool - he made After and Razzle - nice

1h agoHN ↗

If Jev is fundamentally trained using RLCD while you’re building on a Qwen model that was trained using RLHF, how can the resulting model be considered Jev-like?

8m agoHN ↗

for some reason this is really funny to me. it's like the "black museum" black mirror episode where a consciousness in a toy animal can only communicate using predefined responses

31m agoHN ↗

All the people that are just writing an Jev-like API on top of a normal LLM are missing the point. What makes Jev special is the training data; it's how it's trained. The architecture is probably nothing special. Just a text encoder with parallel prediction branches.

I have tried many of these open-source Jev-like models on some linguistic tasks and they are so bad compared to Jev.