Hacker News

Top stories

Live mirror
30 storiesupdated just nowView source snapshot
  1. Ollaya – Ollama for open-source, Jev-style decision models (ollaya.dev)
    50comments
  2. Show HN: Jev Plays Pokémon Red (jev-pokemon.vercel.app)
    36comments
  3. Platform-independent SIMD in Go (go.dev)
    122comments
  4. Bug: Border radius has infected VSCode editor (github.com/microsoft)
    15comments
  5. First Principles Thinking (sunilsadasivan.com)
    75comments
  6. Git-bug: Distributed, offline-first bug tracker embedded in Git (github.com/git-bug)
    90comments
  7. U.S. appeals court upholds designation of Anthropic as supply chain risk (cnbc.com)
    495comments
  8. Alan Kay: Shannon gave us a way of dealing with noisy channels [video] (youtube.com)
    17comments
  9. Advice to a Beginning Graduate Student (2001) (cmu.edu)
    10comments
  10. Show HN: Make math automatic with Mathy (gmays.com)
    4comments
  11. Too AI; Didn't Read (tai-dr.com)
    1comments
  12. Pentium II at 600Mhz with Voodoo 3 Emulated on 86Box with M6 Mac Mini (nyaa.sh)
    108comments
  13. How video games inspire great UX (2019) (jenson.org)
    6comments
  14. Meta's Muse appears to use an OpenAI model labeled muse-special (mouse.dev)
    29comments
  15. Rising sea destroys homes, erases beaches in California (reuters.com)
    17comments
  16. Ink and Switch interactive homepage (inkandswitch.com)
    25comments
  17. Bwbach, My Guardian Goblin (robertmay.photography)
    6comments
  18. Factorio that you can touch (factorio.com)
    80comments
  19. What happens when you analyze your favorite college football team like the CIA? (cultivatelabs.com)
    9comments
  20. Amiga Screens: A Primer (datagubbe.se)
    30comments
  21. Letterboxd Is Up for Sale, and A24, Sony and the New York Times Are Bidding (worldofreel.com)
    13comments
  22. Show HN: Whiteboard (YC W26) – An open-source IDE for thoughtful software design (github.com/devdotfast)
    128comments
  23. A History of the Chiming Machines at Gloucester's Cathedral and Churches (2017) [pdf] (bgas.org.uk)
    1comments
  24. What About Rails? (jardo.dev)
    182comments
  25. Typst makes big strides (lwn.net)
    9comments
  26. Show HN: Doom or Bloom, map your AI worldview (doom-or-bloom.com)
    30comments
  27. Why is the liver so weirdly regenerative? (dynomight.substack.com)
    275comments
  28. CVE-2025-13032: Entering and Breaking the Avast Antivirus Sandbox Part 2 (safateam.com)
    27comments
  29. Boards of Casio (ambionix.com)
    33comments
  30. Astronomer watches Starlink satellites sinking to build a 'planetary barometer' (theregister.com)
    10comments

Ollaya – Ollama for open-source, Jev-style decision models

189 pointsby 2h agoollaya.dev
48 comments
2h agoHN ↗

Are there many models that are comparable to Jev for generic decision making?

Smarter move if you have an eval set is to just train a classifier and call it a day.

1h agoHN ↗

<<<"i was curious to see if i could train a competitive Jev-like model completely autonomously with a swarm of agents using our internal system."

Bro is writing off the H200 lol

On a sidenote I really can't stand the term "swarm" and definately plays into AI doomerism.

2h agoHN ↗

The link rgbrgb posted is a good overview. The best open ones are close to Jev now, but they're big models. And I agree, if you have an eval set for a fixed task, a trained classifier is the better choice.

2h agoHN ↗

Cool... but this does seem undermined by the fact that Ollama can add support for decision models at any time.

2h agoHN ↗

Fair, and I'd be happy if they did. Ollaya uses the same API as Jev, so your code isn't tied to it either way

2h agoHN ↗

and that ollama is go-llama and not rust, so it's not really the ollama of anything

2h agoHN ↗

great project for empowering open-source alternatives.

2h agoHN ↗

open-source is the only way for safe AI development. whoever doesn’t share the weights/code will lag behind.

2h agoHN ↗

Run decision models locally.

example is a text classification task instead of a decision

2h agoHN ↗

Fair point, that example is basically classification. I'll change it to something that looks more like a real decision.

1h agoHN ↗

text classification is equivalente to decision. This is exactly the same thing Jev does.

1h agoHN ↗

Their marketing language is misleading. They must still use some transformer language model backbone to encode the text input (BERT or decoder-only LLM). The biggest difference is the output, instead of auto-regressively generating tokens, they produce probabilities over a bounded set of decisions (more flexible classification).

1h agoHN ↗

It is not. In a benchmark with actual decisions - navigation, traffic, waypoints - laya does only slightly better than a small classifier.

1h agoHN ↗

If it has four legs, a tail and barks why not call it a dog?

1h agoHN ↗

Because this specific dog only barks in structured text

1h agoHN ↗

"Decision model" is just marketing jargon.

decision model = classifier

system one model = small non-reasoning LLM

noul = boolean

confidence = f(probabilities)

It's sad to see how gullible engineers are today.

30m agoHN ↗

how gullible ... today

that laya is even a thing is further evidence, people took that author at face value, the paper contents are incomplete and describe something that does not sound like Jev at all

this was the period of arxiv history that led to the new vouching system, laya author contributed to that imo

2h agoHN ↗

I am fairly confident if Jev-style decision models are seen as prominent (which, they seem to be), Ollama will support them. Surprised the team hasn't implemented this already.

2h agoHN ↗

Sounds good on latency but how is its actual decision quality vs. Jev?

2h agoHN ↗

Depends on the model. The small ones I support today are well below Jev on harder queries, but fine for simple, well-defined questions. The open models that get close to Jev are bigger, and I'm adding support for those next.

2h agoHN ↗

Has anyone actually seen better or the same results with Laya compared to Jev? From my experience, Laya performs significantly worse. It's less confident and often makes wrong decisions with more complex queries.

2h agoHN ↗

Developer here. You're right, Laya is a lot weaker than Jev, especially on harder queries. It's a small model, so it's fast, but that's the trade-off. The open models that get close to Jev are much bigger, and running those is what I'm working on next.

1h agoHN ↗

What are the models? I am super curious in these as well

28m agoHN ↗

Probably Kev and/or the decider models. Kev is trained on one of the 4B qwen models, similar for decider but it ranges from 0.8B through to the 35B-A3B model so far I believe.

2h agoHN ↗

Yes. JEV generalizes better because they probably have an enormous corpus and trained on it for a long time. Laya's out of the box model is much weaker. However, in the age of LLM's it's incredibly easy and cheap to generate large datasets to fine tune laya for your task, and the training loop is pretty quick and cheap too.

It's so easy that I question why I would ever pay for JEV when eventually I'll have done enough random things that I will also have a large corpus and likely a general model as well.

1h agoHN ↗

Isn't the point of Jev that it generalises better?

It's a fast classifier you can use out-the-box, ~1.5bn tokens is about $40 (I've been hammering it)

It just works ... a whole bunch of low-level/low-importance workflow stuff that was getting farmed out to small/fast LLM models now has a competitive alternative ... and bits that hadn't even been considered to go into some external descision/classifier service can be tested/deployed at ~$0.00003/req

I don't get this wall of negativity on it, it's genuinely innovative/useful tech ... would expect HN to be more positive, regardless of whether it's the absolute best execution

1h agoHN ↗

It really does just work. And it works so well I already integrated it into my product. Saves me about 75% of costs for the section its working in, which isn't a small amount. I see a lot of negativity and I don't really get it either. Its so cheap and so fast, why not give it a try?

56m agoHN ↗

I think it’s the infamous Dropbox reaction - anyone can wrap an FTP server, where the innovation?

Starting from a business POV one should inflate terminology, hack together an MVP, and see if the market demands it before doing hardcore R&D.

But starting from technical/craftsman POV all you see is a hack and a lot of big words, so it’s easy to become jaded.

1h agoHN ↗

I didn't see any negativity in the post you replied to.

I think the point being made is that Jev is great but it has no competitive moat, and open source versions will very soon catch up if their secret sauce is just synthetic data.

(Whether or not that is true, I don't know.)

1h agoHN ↗

I think your point is valid but many are annoyed that it is presented as groundbreaking, revolutionary, novel frontier tech when it is a known classification system. It’s the hype that feels undeserved. Honestly it was one of the best marketing campaigns I’ve seen.

1h agoHN ↗

Nothing yet. Unfortunately it sometimes feels like our industry has been overrun by grifters and chancers.

I’m sure this has been a gradual and long decline. Maybe it even started with the dot com boom and accelerated with crypto. With AI it seems to have got worse.

1h agoHN ↗

I've been following jevbench twice a day for the past week and that's been a lot of fun. Latest update:

Rank System Score Public / sealed accuracy Evidence

1 decider-4b v2 64.13 83.5% / 34.7% Evaluator-run, offline

2 Jev 1.13 63.29 86.6% / 36.7% Evaluator-run API

3 JevK5 v0.2 62.04 85.3% / 33.1% Evaluator-run

4 Cygnet 12B 61.76 87.9% / 33.8% Evaluator-run, offline

5 Hopper 59.43 82.3% / 34.1% Evaluator-run

28 Kev 4B 36.14 66.2% / 22.4% Evaluator-run

41 Laya 421M 30.25 58.4% / 30.8% Evaluator-run

https://benchmarkheaven.com/jev-models

55m agoHN ↗

What the best way to see how a homegrown version compares?

1h agoHN ↗

one day, perhaps people will click through to the laya author's arxiv paper content and the why may become clearer, you won't have to read it, a skim will suffice

1h agoHN ↗

It would be really cool to have LLMs and System One in a single tool - in this case, if Ollama implemented it.

1h agoHN ↗

I have also tried this and its really awesome

1h agoHN ↗

Hey Claude, make ollama for Jev like models. Make no mistakes /s

1h agoHN ↗

Guys I have a real q, what is the difference between an instruct based re-ranker and laya/jev I just don't see it.

Edit: One is that jev/laya are tuned to have better probabilities, but a reranker can be fine tuned to do that as well. And jev/laya use RLCD?

42m agoHN ↗

Calibrated probability across multi task with zero shot I guess. A reranker is single task and tuning it make it even more narrow. And I guess some piping to make multiclass efficient since you cannot mask logprob for independent questions in the same output space without throwing calibration away.

36m agoHN ↗

difference between an instruct based re-ranker and laya/jev I just don't see it

Main difference is that laya/jev/et-al give you a zero-shot classifier that requires no training. You can prompt engineer your way to a quick fairly reliable cheap enough decision engine that you can use to iterate quickly (by prompt engineering).

Right now a lot of people are doing this with LLMs and it's too slow and expensive.

Imo the right iterative approach to productionizing these systems is something like:

    1. Build it with an LLM. Iterate on the prompt
    2. Start building a real-world dataset
    3. When the prompt works, turn it into a clear rubric for Jev or similar
    4. Keep iterating until desired accuracy achieved
    5. Use the real-world evals you've built to train a custom classifier fine-tuned to your needs

You now have a system that has produced useful results in production from the very beginning and by the end it's a reliable super cheap classifier that can make thousands of decisions per second.

59m agoHN ↗

Does anyone know what laya multi lang is faster than laya en? I would have thought focusing on a single language would be faster.