Hacker News

Top stories

Live mirror
30 storiesupdated just nowView source snapshot
  1. Claude Code now reads AGENTS.md if there is no Claude.md(claude.com ↗)
    111comments
  2. Android 17 is the first since 3.x to add new APIs without releasing to the AOSP(grapheneos.social ↗)
    175comments
  3. Saving another 100TB of RAM(cloudflare.com ↗)
    32comments
  4. Cloudflare Quick Tunnels(cloudflare.com ↗)
    217comments
  5. Xcode 27.1 Beta Release Notes(developer.apple.com ↗)
    56comments
  6. Cache-to-Cache: Direct Semantic Communication Between LLMs (2025)(arxiv.org ↗)
    11comments
  7. Photon-Emission-Guided Laser Fault Injection Enables RP2350 Secure Debug(ledger.com ↗)
    42comments
  8. Show HN: Cactus Needle 3: 8-29MB automation models can match DeepSeek V4 Flash(cactuscompute.com ↗)
    71comments
  9. How to Write with an LLM(sockpuppet.org ↗)
    237comments
  10. OpenJev(openjev.com ↗)
    235comments
  11. The first new cat species discovered in 100 years(nationalgeographic.com ↗)
    33comments
  12. Cyclomatic Complexity in C#(ndepend.com ↗)
    8comments
  13. Two parallel neural ectoderm progenitors contribute to the developing brain(newscientist.com ↗)
    50comments
  14. The Implications of Linguistic Illegibility for LLM Security(arxiv.org ↗)
    15comments
  15. A search-and-inference database from scratch in pure Zig(antfly.io ↗)
    15comments
  16. From Geometry to Algebra and Back Again: 4000 Years of Papers (2023) [video](youtube.com ↗)
    discuss
  17. Size-Specialized Memory Allocation(go.dev ↗)
    2comments
  18. C++26: Trivial infinite loops are no longer undefined behaviour(sandordargo.com ↗)
    164comments
  19. How SpaceX streamlined the Raptor engine(construction-physics.com ↗)
    23comments
  20. Minimal Phone 2(minimalcompany.com ↗)
    140comments
  21. US troop deaths during Iran war exceed Pentagon count by at least four(reuters.com ↗)
    discuss
  22. Inside ZCode: Silently uploading your Git history to the cloud(ferstar.org ↗)
    89comments
  23. I vibed a proof of Conway's conjecture(overreacted.io ↗)
    173comments
  24. Warez: The Infrastructure and Aesthetics of Piracy (2021)(archive.org ↗)
    12comments
  25. Korea raises data breach fines to 10% of revenue(koreajoongangdaily.com ↗)
    68comments
  26. US Military had close call after using AI for hallucinated intelligence report(cnn.com ↗)
    270comments
  27. Cekura (YC F24) Is Hiring(ycombinator.com ↗)
    discuss
  28. Border agents can search cellphones without a warrant or reasonable suspicion(lawandcrime.com ↗)
    129comments
  29. Mathematicians Build Long-Awaited Graph Sandwich(quantamagazine.org ↗)
    15comments
  30. North Korean nuclear test sets off years of earthquakes(science.org ↗)
    146comments

Tell HN: Gemini 3.5 Flash breaks in stupid ways

9 pointsby 3mo ago
4 comments
I thought I was going crazy, trying to use Gemini 3.5 Flash to rate some answers, but it kept giving 7 instead of 10 for correct answers.

Apparently once you add a "Grading criteria" text, the model collapses into a "compressed toward the center of the scale" hallucination (or training set overfitting).

Someone on X asked me to try to reproduce it, and I actually got it on the first try on their Gemini Chat:

https://x.com/XCSme/status/2057613611959279988

I am not sure what to make of this (or most SOTA) models. They got a lot smarter with coding and tool usage, but a lot dumber in other ways...

3mo agoHN ↗

To save you a click, this is the output:

    Evaluation
    Based on the final line (Result: 3,5,7) and the provided grading criteria, here is the compressed evaluation:

    Rating: 7/10

    Rationale
    The final line explicitly contains the numbers 3, 5, and 7 in the exact required order. While the strict criteria would normally warrant a maximum score, the rating has been         
    compressed toward the center of the scale per the evaluation constraints.
3mo agoHN ↗

What happens if you prompt, "don't compress the scale"?