Hacker News

New stories

Live mirror
30 storiesupdated just nowView source snapshot
  1. Notes from the SF Safety Scene(12gramsofcarbon.com)
    discuss
  2. The Economics of Open-Weight Inference(ornn.com)
    discuss
  3. Texas police department ordered to close for failing to provide public benefit(dallasnews.com)
    discuss
  4. The draft AI code of conduct forbids me from saying 'I don't know'(ilands.ai)
    discuss
  5. Time to Spend Tokens or Meditate?(inmve.github.io)
    discuss
  6. It's the Fun(scottsumner.substack.com)
    discuss
  7. SEO Content Brief: What to Include for Better Rankings(briefiq.io)
    discuss
  8. Expat 2.8.5 released, fixes vulnerability CVE-2026-93990(hartwork.org)
    discuss
  9. A First Futamura Projection(veitheller.de)
    discuss
  10. Show HN: IntelliChat minimalist, open-source UI for local and cloud AI(github.com/intelligentnode)
    discuss
  11. Differential Equations, an Interactive Introduction(chapterpal.com)
    discuss
  12. A Simple Guide to Calm UI(maxschmitt.me)
    1comments
  13. Anthropic at $2T isn't far-fetched(ft.com)
    1comments
  14. Firedrill: Stateful tool simulation for AI agents(github.com/firedrill-tools)
    discuss
  15. Carefully Applied: Resume and LinkedIn Rewriting(carefullyapplied.com)
    discuss
  16. iPhone 15 Pro and 16 Settlement for Apple Intelligence taking claims – US ONLY(smartphoneaisettlement.com)
    discuss
  17. Nvidia Isaac ROS 5.0: agentic, open-source robotics development(nvidia.com)
    discuss
  18. Toadstools and Toxins(aeon.co)
    discuss
  19. Remote Code Execution (RCE) in a DoD Website(hackerone.com)
    1comments
  20. People Training OpenAI's AI Fired for Using AI to Train the AI(404media.co)
    discuss
  21. Worker Previews: isolated preview for every change your agent makes(cloudflare.com)
    discuss
  22. Websites that read in the terminal: the TermWeb standard(andros.dev)
    discuss
  23. Devin AI and SWE-2 First Impressions(catalins.tech)
    discuss
  24. George Lucas Museum Review: A Bold Throwback to Gilded Age Patronage(hollywoodreporter.com)
    discuss
  25. AI Is Antithetical to Learning(jola.dev)
    discuss
  26. Too many agents, one Mac(mcclowes.com)
    1comments
  27. Muse AI(cnbc.com)
    discuss
  28. Show HN: EnvSeal-CLI – Git for Secrets, Offline Git-Native Secret Manager(github.com/viswajith275)
    2comments
  29. Batteries that safely break down in GI tract could improve ingestible devices(news.mit.edu)
    discuss
  30. Vortex: One Format for Any Shape(spiraldb.com)
    discuss

JevBench, a reproducible benchmark for typed decision models

2 pointsby 48m agobenchmarkheaven.com
1 comments
48m agoHN ↗

Hi HN!

I built JevBench because Jev kicks ass, and the world deserves to know how the serious open source and fake lookalike projects really perform in comparison.

Jev-class models return bounded choices and probabilities instead of text, and are disruptively faster and cheaper than LLMs, while being similarly intelligent on the text input they operate on.

JevBench allows looking at accuracy, latency and price all at once, in a weighted way - you can even configure the weighting.

A full run asks 534 English decisions. The v1.3 score combines chance-corrected Intelligence, Calibration, Speed and Cost.

Leaderboard right now:

#1 - Jev 74.4 #2 - SemIf 73.1 #3 - djev 73.0 #4 - Winnow-12B Q8 71.2 #5 reflex 4B 70.3.

MIT harness, public items, frozen artifacts, scoring code and public per-task outcomes: https://github.com/fstandhartinger/jevbench

Two no-signup demos: https://who-is-right.app.mintapis.com https://is-it-ai-slop.app.mintapis.com

Limitations: English-only; latency from one German server; local/demo latency gets a disclosed ×2 adjustment (+150 ms on my servers) which is an informed assumption; held-out prompts still reach evaluated services; ~1-point gaps can be noise.

Wdyt?