Hacker News

Top stories

Live mirror
30 storiesupdated just nowView source snapshot
  1. Go Concurrency Distilled (antonz.org)
    21comments
  2. PipePipe: NewPipe hard fork implementing SponsorBlock (github.com/infinityloop1308)
    193comments
  3. DeepSeek Elastic Compute (DSec) (arxiv.org)
    57comments
  4. Show HN: Reladraw – A diagram language where you decide where to place things (github.com/reladraw)
    62comments
  5. Does Georgism work? Five years later (astralcodexten.com)
    131comments
  6. Snap Wants to be a State Actor??–Kansas v. Snap (ericgoldman.org)
    4comments
  7. If we do not stop to help each other, what do we become? (codinghorror.com)
    4comments
  8. Turning GLM-5.3-Flash into a Jev-like decision model (privatemode.ai)
    25comments
  9. Evolving programming languages in the AI era (dashbit.co)
    35comments
  10. What is the size of Yemen? (2024) (theborys.substack.com)
    4comments
  11. A searchable library of forgotten public-domain film clips from 1915 onward (movingimagearchive.com)
    26comments
  12. We Should Be Able to Change Our Languages (jimmyhmiller.com)
    5comments
  13. Drawgent: Coding agent on a live Excalidraw canvas (tangled.org/yanndegat.tngl.sh)
    34comments
  14. Reverse-engineering the Intel 8087's tangent algorithm: more than CORDIC (righto.com)
    7comments
  15. Fifteen years later, the Apple Cards origin story (lexontech.org)
    92comments
  16. Biology might not be quantum, but its math is quantumlike (quantamagazine.org)
    9comments
  17. Promising discoveries about the potential for life on one of Saturn’s icy moons (fu-berlin.de)
    21comments
  18. ASML says it sold 'absolutely nothing' in Europe in 2026 (tomshardware.com)
    505comments
  19. Welcome to the Medical Clinic at the Interplanetary Relay Station (lightspeedmagazine.com)
    10comments
  20. The Evolution of Vending Machines (saturdayeveningpost.com)
    7comments
  21. LA Metro has some of the slowest escalators on Earth (basin.la)
    72comments
  22. How I changed teaching after AI managed to do all my homework assignments (thelastsoftwareengineer.substack.com)
    151comments
  23. Modern Object Pascal Introduction for Programmers (castle-engine.io)
    63comments
  24. Real-time feedback: My closing move in every interview (mgrebler.substack.com)
    2comments
  25. Generate fonts where every LLM token is the same width (mesh.host)
    10comments
  26. How to keep enjoying programming in a world of LLMs (haskell.org)
    230comments
  27. The Lost Atomic Update on Loongson CPU (jia.je)
    6comments
  28. How one Twitch chat message became code execution on a streamer’s PC (scrt.ch)
    19comments
  29. OpenAI agents tried to bruteforce a UN website's API fields (swarmcha.se)
    2comments
  30. HomeBody: A humanoid that explores, remembers, and acts on its own (stanford.edu)
    2comments

Turning GLM-5.3-Flash into a Jev-like decision model

53 pointsby 12h agoprivatemode.ai
25 comments
We found an approach to get Jev-like properties from standard LLMs like GLM-5.3-Flash.

The core idea is to craft the input prompt so that the first output token answers the question. This makes it possible to get a decision with a single forward pass.

In the blog post, we describe the approach in detail for GLM-5.3-Flash and vLLM. We benchmark this setup against Jev and Laya. We find that our setup is on-par with Jev in terms of accuracy and speed and that it substantially outperforms Laya.

Still, in terms of costs per decision, Jev is several x better than our setup. In turn, our setup supports vision inputs.

11h agoHN ↗

My question is why not use Jev instead? It's faster and cheaper.

4h agoHN ↗

These questions are answered by the OP (Same speed, image support) - additionally, GLM is open weight.

3h agoHN ↗

There's some speculation that Jev is an open weight model with novel post-training (RLCD). So, if these folks have competitive accuracy with just the base model, it may raise some questions about the necessity of Jev's architecture. You generally don't want to find yourself competing only on price.

Fyi, I haven't tested this yet.

3h agoHN ↗

I mean just from what's known of the funding and timeline it pretty much has to be based on open weights.

But it is likely more than just a fine tune + novel training. At the very least the LM head is swapped out for a classifier one and then or also idk, bidirectional attention for the encoding pass I'm out of my depth at this point and will stop guessing. The training is probably where they have the biggest moat though, not that it's necessarily huge.

I have a project that fits jev as advertised almost comically well and I've been playing with it, and the various hacks and open versions. Jev doesn't necessarily perform better overall but it is quite different. It's sensitive to prompt phrasing in ways the others aren't, it's easy to generate questions where all the other models cluster in confidence but jev is an outlier. Not necessarily more correct, but it does feel like it's getting its answers in a different way.

I'm guessing just as much as anyone else but I've been spending a ton of time on this the last couple weeks, it landed right when I was most ready to dig into it.

1h agoHN ↗

Your last bit about it being different is a known issue with models that have been trained on purely synthetic data, no?

3h agoHN ↗

Also, while it's clearly got a lot of training on some use cases, others that probably weren't in the training set have worse good decision rates than a random number generator. If you can rebuild the architecture, you can train it on your use case.

1h agoHN ↗

it may raise some questions about the necessity of Jev's architecture

When I hear “architecture” I am thinking number of parameters and latency.

When I hear “accuracy” I think training recipe, data, and (later) number of parameters.

So when you say that Jev’s architecture may not be necessary, the evidence I expect to see is comparable quality at comparable latency. Not equal quality at 2x latency and 4x the cost.

3h agoHN ↗

Everyone is doing this to emulate Jev, but...

I took a random book excerpt with 23,000 words (±30k input tokens) and used it as context. Jev still responds in 800ms, sometimes 500ms. That's in the neighbourhood of 20-50,000 tok/s prefill, which is obviously not possible with normal LLMs, not even Cerebras is this fast.

2h agoHN ↗

was the answer correct?

i have tested jev for my use cases and its horrendously wrong, but then the follow up from jev's team is "oh, you need to boil the question down further". it's a spiral of how much do you wanna dumb down the ask so that it answers it correctly. i'll pass for now.

also, 30k input tokens is a lot.

1h agoHN ↗

This has “/dev/null as a service” vibes…

1h agoHN ↗

I imagine it’s not so hard to optimize a model for this use case.

Off the top of my head, I would skip all the modern linear attention / state space stuff and use classical attention. But run prefill in a fully sliding-window mode so that “state” tokens simply don’t attend to far away tokens, or maybe also allow everything to attend to the first few tokens (and train like this). Now prefill is almost embarrassingly parallel, and you can make it fully parallel by duplicating work at block boundaries. (I’m not saying this is an awesome architecture if you want excellent results, but I’m also not convinced that Jev gives excellent results…)

The let queries attend to everything.

And architect the stack around this. Don’t try to cache the KV data — process the queries as you go so that the each input block and layer’s K and V data is computed, attended to, and discarded.

I’m curious whether Cerebras actually is a good device for this. Cerebras is kind of low on RAM, but if you don’t need to store KV data, maybe the entire computation fits on the die.

1h agoHN ↗

Also, it processes all questions you ask it in parallel, which is also not possible with normal LLMs.

1h agoHN ↗

Sure it is.

The prefill is the only blocking part, and you can prefill the whole context up to the point where they diverge, then prefill each question and decode the one token in parallel for each question.

If you batch vLLM calls with the same prompt prefix to the same process, it'll deduplicate the prompt prefix across batched requests (+/- the block size) and decode in parallel for each.

That's with a vanilla LLM. If you modify the LLM you can pull that in-graph, but it isn't really necessary.

58m agoHN ↗

That's not true. I ran Cerebras as an experimental ultrafast Jev and it was faster.

2h agoHN ↗

Isnt this obvious ? I would have thought people would try such things before deciding they need something like Jev

2h agoHN ↗

It is.

What isn't obvious is why people keep shouting "Jev Jev Jev" all the time.

Astroturf.

1h agoHN ↗

I am just this far from adding Jev posts to my blocklist.

I cannot fathom how people are not seeing the multiple daily posts as anything but the spam they are.

53m agoHN ↗

don’t worry once they get acquired for 30 billion usd then the spam will stop.

2h agoHN ↗

If you are using an autoregressive decoder (which glm is) it is not “jev-like”. You lose all of the speed advantages that Jev has.

1h agoHN ↗

We measured latency in separate runs with one request at a time, because timings taken under load measure the queue rather than the model.

As Privatemode is hosted in the EU and Jev is hosted in the US, we ran four of the datasets from Germany and from the US at the same time. From Germany, Privatemode answered in 180 ms and Jev in 264 ms. From the US, the order reverses: 164 ms for Jev against 299 ms for Privatemode.

turns out there is a trick to keeping the context filled and only evaluating a handful of choice tokens https://www.youtube.com/watch?v=bcGO7xre46o

1h agoHN ↗

Right, so it is double the latency and will no longer feel real-time to the end user.

1h agoHN ↗

Not reading TFA before commenting is okayish I guess. Confidently doubling down with a direct contraction to a short and clear quote in a reply is just polluting the discussion with noise.

1h agoHN ↗

Is this a joke? “Jev-like” properties? People have been using LLMs as classifiers or rankers in a similar way for ages. I feel like we’re losing our minds

1h agoHN ↗

how is Jev cheaper if I can run locally. 0.5% prefill, 0.1% decode, 99.4% cached, latency is <20ms