Just curious, where has this term 'noul' come from for yes/no ansers?
/a bit more digging and..
A Noul performs a Bernoulli trial—an experiment with exactly two outcomes (yes or no)—but instead of picking one, it returns the calibrated probability (ranging from 0.0 to 1.0) that the statement is true.
Not quite.. that boolean is about whether the voltage exceeds some threshold. It's not about how close the voltage is to the circuit's maximum possible threshold, or how much it exceeds the threshold.
Seems like this is the way, a hybrid approach where some of the pipeline will be jev like and some traditional LLM depending on the nature of the work.
My bots are all named Jeeves lol. I have a CLI tool I use that connects up to a LLM I made and I call it Jeeves too ...so funny. I really didn't use Jeeves all that much I tended to use...I think it was called Web crawler pre-google era
interesting bench list, what about benchmark against smaller or bigger models? 9B looks too huge for small like laya, and too small for llm-level decisions.
Just curious, where has this term 'noul' come from for yes/no ansers?
/a bit more digging and..
A Noul performs a Bernoulli trial—an experiment with exactly two outcomes (yes or no)—but instead of picking one, it returns the calibrated probability (ranging from 0.0 to 1.0) that the statement is true.
I hate it :)
Bernoulli
If you want to get super pedantic about what’s happening in a transistor every digital Boolean is actually this
Not quite.. that boolean is about whether the voltage exceeds some threshold. It's not about how close the voltage is to the circuit's maximum possible threshold, or how much it exceeds the threshold.
In an analog circuit, maybe.
The whole "no hallucinations" premise is based on that.
Like, yeah, you don't hallucinate, but only because you force the user to decide in the end.
Wait til you hear about how digital circuits work at die level
That seems like the only possible way to eliminate hallucination, short of a model that is never wrong.
force the user to decide in the end
And that's...bad?
I like it. It is short and distinct which is a good fit for a primitive. It describes its fundamental meaning and draws a connotation with Boolean.
In Bayesian statistics that’s called credence. Weird that they felt the need to invent a new term.
Are there "good" Open source Decision models built on Gemma-4 and also trainiable on own data?
see: https://news.ycombinator.com/item?id=49883844
Jeeves, that's a name I haven't heard in a long time...
If Jeeves returned as an AI chat bot it would be the most brilliant resurgence of nostalgia
jeeves is currently the name of my local hosted assistant, in its context there are rules that tell it to behave like good old jeeves.
soon I'll make sure that my home assistant pod answers to "Hey jeeves"
Personally I’m happy that after a 30 year effort and hundreds of billions spent, AskJeeves finally works as intended.
We locked him in the basement with Clippy, Bob and BonziBuddy. Who opened the damn basement door???
…a long time.
This isn't really surprising. LLM reasoning and before that, chain of thought prompting are essentially forms of test-time compute scaling.
Seems like this is the way, a hybrid approach where some of the pipeline will be jev like and some traditional LLM depending on the nature of the work.
So "askjeeves" has been resurrected?
Ask Jeeves - Only took us 30 years to come full circle.
this was the comment i came here for
My bots are all named Jeeves lol. I have a CLI tool I use that connects up to a LLM I made and I call it Jeeves too ...so funny. I really didn't use Jeeves all that much I tended to use...I think it was called Web crawler pre-google era
Super dope. If it would ship as prod ready code supporting mps as well that would be even doper.
But funny that jev is getting its lunch eaten apparently in under two weeks?
It’s ok, one week of AI hype is now enough to close a billion-dollar term sheet with VCs.
interesting bench list, what about benchmark against smaller or bigger models? 9B looks too huge for small like laya, and too small for llm-level decisions.
See if it can beat Jev's Pokémon benchmark
what benchmark this one or it's just fun?
what about benchmark against smaller or bigger models? 9B looks too small for llm-level decisions.
I'm surprised we haven't seen a "Jehovah" yet.
With the amount of talk about "inventing god", I'm surprised too.
What is the point of this, if it is p90 17 seconds? Might as well use an LLM. The beauty of Jev is that it is dirt cheap and insanely fast.
I would hold your horses to paint it as dirt cheap.. In my cases for spam detection Luna was 20% cheaper due to prompt caching, although not as fast.
Jev ought to offer a flex mode that uses their spare capacity for a discount.
Jev-like models give calibrated decision probabilities, but at low accuracy.
So why didn't they show both??