- 39comments
- 396comments
- 9comments
- 108comments
- 64comments
- 236comments
- 15comments
- 333comments
- 26comments
- 259comments
- 263comments
- 51comments
- 323comments
- 8comments
- 25comments
- 1comments
- 78comments
- 2comments
- 2comments
- 29comments
- 66comments
- 20comments
- 62comments
- 38comments
- 59comments
- 39comments
- 78comments
- 483comments
- 23comments
- 198comments
About jev being a black box and the potential for bias, I think it boils down to what questions you are asking the model.
Broad questions like Is this resume good / score this city will ofcourse be biased but I think jev encourages more granular focused questions like Score this candidates Python experience / Rate this city for its food which then allows you to introduce your own biases in which questions you ask and how you combine their answers.
In this way I think jev like models can be easier to reason about for critical decisions.
I’m not sure I understand the hype around this model. Isn’t this just an llm with a chat template, with the options prefix cached?
Then the llm is constrained to a few special tokens indicating the possibilities? e.g. <option1> <option2>
To get calibrated probabilities sounds like a very good feature, if they are indeed well calibrated.
And in my experiments even Qwen 3.8 has a hard time to consistenly conform to a schema, requiring retries, JSON cleanup etc, so to have a model of similar quality (SemIf et al) that simply cannot deviate from the schema by construction could be very helpful.
But I still need to experiment with either Jev/SemIf myself.
Is it 100% deterministic?
I think "Black boxes are back in fashion" is missing the point. I think LLMs are still largely black boxes, and I don't think chain of thought is representative of any degree of inner machination. Asking it questions to justify itself is at best a facsimile, and for the most part it's useful, but it's fundamentally a facsimile.
Where I understand Jev to be a significant jump is that afaik the confidence scoring is actually derived from the normalised probabilities, and not a continuation in a chain of prediction masquerading as "confidence."