Hacker News

New stories

Live mirror
30 storiesupdated just nowView source snapshot
  1. Blocklist Whitelist Anything on YouTube or YT Kids
    discuss
  2. A lesbian bar required KN-95 masks. When this changed, all hell broke loose(bostonglobe.com)
    discuss
  3. How Zambia Approved an HIV Drug Quickly(asteriskmag.com)
    discuss
  4. Silicon solar cells could cut satellite power costs by up to 90%(surrey.ac.uk)
    discuss
  5. Show HN: Run Linux kernel drivers in userspace on Neptune OS, NT-like OS on seL4
    discuss
  6. Show HN: Data structure that lets different parts of math work together(github.com/tmilovan)
    discuss
  7. Show HN: OnDevice Tools – 64 browser utilities that never upload your data(ondevice-tools.org)
    discuss
  8. What Is an AI Software Factory? Lessons from 3 Client Deployments(camplight.net)
    discuss
  9. Sen. Bernie Sanders unveils bill to ban artificial superintelligence(apnews.com)
    discuss
  10. Break Your Network on Purpose(github.com/arjunshajitech)
    1comments
  11. Build a Browser Extension with Jev(medium.com/towards-artificial-intelligen...)
    discuss
  12. PostMCP – Turn any OpenAPI spec into a context-optimized MCP server(github.com/braveram)
    discuss
  13. Trained KV cache bank turns any LLM into Jev like Model(ts.net)
    2comments
  14. My German Exit Tax: 13 Tax Advisors, 1.5 Years, 18.7k€ to leave the country(eidel.io)
    discuss
  15. The Alien Signal That Looked Intelligent [video](youtube.com)
    discuss
  16. Zipf's law(wikipedia.org)
    discuss
  17. Show HN: Codex Reset – A tracker for public Codex reset signals(codexreset.live)
    discuss
  18. Claude Microfactory(claude.ai)
    1comments
  19. London buses in decline? Why public transport in global cities is slowing(theguardian.com)
    discuss
  20. Longtime SUSE staff asked if they'd opt for 'voluntary separation'(theregister.com)
    1comments
  21. Nemotron-H: A Family of Accurate, Efficient Hybrid Mamba-Transformer Models(nvidia.com)
    discuss
  22. Show HN: What's Next – make Claude Code respond with choice, not prose(github.com/kvitapp)
    discuss
  23. Enables Nvidia DLSS Frame Generation (DLSS-G) on RTX 30-series and RTX 20-series(github.com/sdli1995)
    discuss
  24. Show HN: ToolRoots, local alternatives to the tools you use(toolroots.com)
    discuss
  25. Qwen-Audio-3.1-TTS(aliyuncs.com)
    discuss
  26. Tokens Too Cheap to Meter(jyn.dev)
    discuss
  27. MNT Station – A modular, open hardware desktop computer and server(crowdsupply.com)
    discuss
  28. Directus The back end for your whole team(directus.com)
    discuss
  29. Next version of OpenClaw uses a decision model to decide between steer or queue(twitter.com/steipete)
    discuss
  30. Show HN: A running app that turns your runs to a garden(apps.apple.com)
    discuss

Jev in 25 Lines of Python

152 pointsby 2h agonobodywho.ai
49 comments
2h agoHN ↗

Latency and compute comparison needed.

1h agoHN ↗

Is benchmarking Jev still a ToS violation?

1h agoHN ↗

Was it? That would make it unusable in any corporate setting.

2h agoHN ↗

Whilst I do like reading these things for technical know how, I can sympathise with the creator of jev who now presumably has to apply an order of magnitude effort to explain why the 100 smaller things done better than this add up to a much better product.

1h agoHN ↗

Replace 'explain' with 'sell'. Don't forget that it's a gold rush. There's no reason to sympathize with corporations in their rush for the slice of the pie.

26m agoHN ↗

You can be unsympathetic to the corporation’s bottom line while being sympathetic to the human beings that had their work trivialized by some cocky blog post.

1h agoHN ↗

IT's the infamous "OneDrive in 10 lines of code (SFTP)"

While technically correct, it's not the same thing

59m agoHN ↗

It's not the same thing but it have advantages Jev don't have like ... being local.

34m agoHN ↗

From what I’ve learned about Jev I feel it’s just a very successful marketing campaign to developers not fully understanding data science (and deep learning). It’s nothing new, been around since 2022? Being local is an extreme advantage lol.

2h agoHN ↗

Beyond the missing latency and compute comparisons that Heaney commenter mentioned, also nothing about its error rate compared to Jev (nor if it even always outputs in a format the app can parse, not sure how solved that is).

But then at the end it says it’s parody. Maybe HN title should say it’s a joke.

1h agoHN ↗

latency and compute comparisons highly depends on your local setup.

you can swith to a better model for lower error rate.

1h agoHN ↗

Which massively slows down the output. Doing this with Qwen 9B already takes you into seconds per answer territory, and Jev is supposedly frontier level intelligence.

2h agoHN ↗

Going directly for the logprobs is always icky when you use a chat model as base, because they are trained to write prose as output. So your "choice" tokens and thus their probabilities might get diluted in whatever else it wanted to say. If you have to do it in the same way as this post, at least add clear system instructions and a carefully worded beginning to the assistant output section of the prompt to lower the chances of it wandering off immediately.

I've found that using structured outputs solves this problem much better. Instead of letting a model generate only "A", "B" or "C" and looking at the probs, have it directly generate "Legitimate", "Spam" or "Phishing" or any other pre-defined option from a set of multi-token sequences. Behind the scenes it boils down to something quite similar, but you're not running into the risk that the model actually wanted to say "A phishing attempt seems likely, so answer (C) is correct.", which would lead "A" to have the highest probability in the first token. You can even use a reasoning budget this way either via inherent reasoning or a free-form part preceding the remaining output structure. You can also have it assign probabilities (either in words or numbers) using more complex output structures, but I would not rely on them much more than the token logprobs (they can still be quite good though).

1h agoHN ↗

Seems like all normal english words could risk the same, so would using short but random strings be even better?

Actually to me it sounds it could be benchmarked if this kind of effect exists in the first place.

1h agoHN ↗

Best option would be reasoning + clear system instructions + constrained output. That is, if you have to use a chat model. Which works well enough to be sure, but hey I haven't tried raising millions of dollars when I did that 3 years ago. But perhaps I was the stupid one.

1h agoHN ↗

Agreed, it's a real issue, but it can probably be vastly reduced by having the schema in the system prompt and by giving the model an expectation of a fixed value: no decent modern would pick a prose ligament over a provided value.

To completely squash the issue, a few cheap LoRa iterations will do the trick just fine.

1h agoHN ↗

Sure, you can fix that in a couple lines. Then a couple more lines for evaluating multiple questions on the same answer in parallel. Then a couple more lines for the confidence score (which is trivial to compute from all we have, but missing regardless). Then a harness to fine-tune an existing model to perform better on this specific task, and a collection of training data to use for that

I think we can all agree that Jev is not rocket science. It's a good idea executed well, with marketing that might have been a tad too bold

1h agoHN ↗

The confidence score is not trivial to compute. That is the whole point of the model. Even if you are using a proper scoring function such as NLL, it is not enough to ensure calibration in deep nets. So you have to do good post training to ensure it. These are all known techniques, but they are far from trivial, especially on large scale datasets.

47m agoHN ↗

Their docs at https://docs.typesafe.ai/confidence state "confidence is a statistic computed from the probability distribution the answer already gives you. TypeSafe computes it for you"

And further down "TypeSafe computes confidence from how the probability is spread across the options. All of it on one option gives 1.0; the more evenly it spreads, the lower the confidence. This demo uses (3 × largest probability − 1) / 2 to approximate confidence for three options."

So while we don't know the exact formula they use, it is just a function over the probabilities

I am open to the argument that this does not work well if you just plug in a qwen model instead of a model that is trained to output more statistically useful token distributions

41m agoHN ↗

I am open to the argument

we agree then, that is the entirety of my argument. Getting a deep net especially one that is anywhere near even SLM size to be calibrated is tough, especially across domains. They claim calibration across a variety of datasets which is interesting.

1h agoHN ↗

In my experience as well using logprobs to try to quantify uncertainty, LLMs are a poor fit. Neural nets in general struggle with 'calibration' --- ie. if a prediction is truly 50/50, neural nets are often prone to predicting overconfidently [0].

I ran some tests using GPT-4 to do some basic classification a couple years ago. On ambiguous options which had to be escalated to a human, the LLM would regularly output something like a 99.8% probability, compared to 99.99% for a correct answer.

0: https://arxiv.org/pdf/1706.04599

1h agoHN ↗

Yes. But even then, the probabilities are not calibrated. In jev/laya, they are (well, relatively anyways).

1h agoHN ↗

That's the approach that daseinlabs/open-jev takes, in contrast to the above, which is what TheoLeeCJ/openjev and ekzhang/openjev-sglang do

https://sgnt.ai/p/jev/

50m agoHN ↗

A fundamental benefit of LLMs over Jev is that you can use test-time compute to improve the accuracy. Jev might eventually evolve to use test-time compute, but the formulation seems to more elusive to me than for LLMs.

47m agoHN ↗

The whole point is the quantified output. If you just ask an LLM to type out its confidence "manually", it'll make up some nonsense. The logprob numbers are more reliable.

I got this technique to work extremely reliably last year. However there were a bunch of caveats: 1) Firstly, you must institute a check that the multiple choice tokens dominate the output distribution. They should sum to 95% or more, ideally 99%, or the LLM is not following instructions properly. This is also the problem with constrained decoding - if the LLM really doesn't want to output a valid answer, the one you extract will not be high quality. 2) You need to ask it multiple times, permuting which option corresponds to which letter, and average the results. LLMs are surprisingly biased towards picking "A", especially if they're otherwise not sure. 3) For the same reason, performance improves if you frame the prompt as if it were the middle of a quiz. "Question 1" carries baggage that "Question 12" doesn't. 4) You must be exceedingly careful with tokenization.

But when all was said and done, I got a general purpose A/B classifier that gave high resolution quantitative output for the cost of a couple dozen tokens ingested and a couple inference passes.

1h agoHN ↗

Now, can you do it in <200ms for 45 questions at once, have 0% malformed output, and any kind of meaningful benchmark? We’ll wait!

1h agoHN ↗

<200ms for 45 questions at once

Considering your own question length: ~120 characters x 45 divided by 4.1 ~= 1317 tokens.

So question processing at 5.5k PP(around the actual PP speed of GPT5.6 Sol) it would take around ~0.24 seconds + the context processing.

Computing the output should be around ~20ms (at 50 tok/s), computing 45 tokens in parallel.

have 0% malformed output

Pretty trivial; only the allowed output is selectable :)

So, I keep repeating myself: Jev was a low-hanging fruit all along; no one cared, and probably no one will in a few weeks?

52m agoHN ↗

You can probably even share context between questions by cleverly manipulating the attention mask.

48m agoHN ↗

Yeah but a lot of developers who didn't even know that this was a possibility now do, and will probably find use cases for it.

55m agoHN ↗

Nothing has malformed output if you coerce it's output into a statically defined set of options

1h agoHN ↗

It's fast.

If you're comparing with something, you need to state 'fast' in relative terms. Jev is definitely fast, and if this Python takes the same time to get a decision then it's also fast. If it's 100* slower than Jev though, you shouldn't be calling it 'fast', because relatively speaking it's really, really slow.

1h agoHN ↗

By design it can't be significantly slower than Jev: the prompt processing (AKA PP) is exactly the same on both and will take most of the time. Then you can process every single "question" in parallel, just predicting one or two tokens (if an answer is ambiguous with a single token) per each question, again in a single batch.

So, fast in the LLM space and comparable with Jev.

1h agoHN ↗

What I don’t understand is, why would you not want “reasoning” in a classifier?

Speed and cost are obvious reasons, but isn’t this a tradeoff?

1h agoHN ↗

not sure if true, but if you look at laya they use BERT type models. If jev is also using a BERT-type model it is autoregressive and therefore can't reason in the way that GPT-type models can. However, you get the advantage of being able to attend in both directions.

1h agoHN ↗

I wonder if this could be a good stepping stone to write a local prompt router to optimise what model get what prompt. I.e. if the prompt is just a lookup, send it to haiku, if it's reasoning, send it to opus and if it's implementation send it to sonnet.

1h agoHN ↗

I was thinking the same. Haven't tried it out.

1h agoHN ↗

Nothing I hate more than bullshit articles claiming X in Y lines of code, only to use libraries abstracting hundreds of thousands of lines of code.

1h agoHN ↗

Should they be writing quicksort in assembly as a first step? I think its legitimate in this case given that Jev is likely using the same tools as the example. Showing how easily the core is created using those tools helps to dispel some of the mystery and hype.

Example why its legit:

I just invented a new "Regression Estimate Validator" aka Rev. It takes hundreds of input dimensions, then outputs an interpretable score. Its very fast and statistically robust. Response: Ok but you could just use `pytorch.nn.Linear(d_in, 1)`? True, it is equivalent, but that's concealing millions of lines of hand-tuned math libs, CUDA, python, and other stuff.

The fact that there are many lines of code underpinning the target functionality doesn't make it any harder to use, and doesn't increase the value of the sales pitch for the "new shiny thing" using those few lines of code.

However, I do sympathize with your frustration that people can just say "its 1 line of code" when that line is "invoke API" which is really millions of lines / databases, etc. as a way to dismiss legitimate work without understanding its implications.

1h agoHN ↗

strong "You can build dropbox quite trivially by getting an FTP account, mounting it locally with curlftpfs, and then using SVN or CVS on the mounted filesystem" vibes

You have built something like jev but not jev (for starters, the output of what you've built will be absolutely worthless, the whole reason Jev is getting so much hype is because the output is good enough)

1h agoHN ↗

Because of masked attention in LLMs, if you put the options before the body (the email to analyze), the transformer already knows what it needs to look for, and can use more tokens to create state to address that specific task (BERT has no mask in the attention, so tokens attend also to next tokens). You could also do a few examples in the system prompt to improve calibration.

Another trick that works is to repeat the question two times: "I'm repeating the task and labels for clarity: ..."

43m agoHN ↗

You can also go beyond Jev. Qwen 3.5 0.8B is fantastic at basic image classification/question answering (including OCR elements) also. Though rather than looking at logits, I get it to output a structured JSON object and it does simple object classification tasks on a Mac at under 500ms a pop (I forget how far, but I think it's like ~250ms) with good accuracy (depending on task).

42m agoHN ↗

What I'm missing here is also type guarantees. I don't think you can do it without token level logic which forces the model to output the tokens from a predefined pool of tokens. A logic like this given some JSON schema is not that difficult to implement. If the LLM must output JSON schema compatible value then you can also add that it doesn't "hallucinate". Which is funny too because just guaranteeing the type does not mean the model does not hallucinate but this is another story.

35m agoHN ↗

Pretty interesting how a simple example like this makes the idea so easy to understand.

33m agoHN ↗

Startup coming out of 2 years of stealth to be reproduced this easily

12m agoHN ↗

Highly suspect of content marketing.

Ends with referring to a product, and saying "this is a parody post", after pretending to make a serious point.