Hacker News

Top stories

Live mirror
30 storiesupdated just nowView source snapshot
  1. Claude Code now reads AGENTS.md if there is no Claude.md(claude.com ↗)
    144comments
  2. Android 17 is the first since 3.x to add new APIs without releasing to the AOSP(grapheneos.social ↗)
    199comments
  3. Saving another 100TB of RAM(cloudflare.com ↗)
    36comments
  4. Cloudflare Quick Tunnels(cloudflare.com ↗)
    231comments
  5. How OpenAI Used Its Own LLMs to Design Its Jalapeño Chip(ieee.org ↗)
    7comments
  6. Xcode 27.1 Beta Release Notes(developer.apple.com ↗)
    62comments
  7. Cache-to-Cache: Direct Semantic Communication Between LLMs (2025)(arxiv.org ↗)
    11comments
  8. Photon-Emission-Guided Laser Fault Injection Enables RP2350 Secure Debug(ledger.com ↗)
    46comments
  9. How to Write with an LLM(sockpuppet.org ↗)
    250comments
  10. Show HN: Cactus Needle 3: 8-29MB automation models can match DeepSeek V4 Flash(cactuscompute.com ↗)
    72comments
  11. US troop deaths during Iran war exceed Pentagon count by at least four(reuters.com ↗)
    42comments
  12. The first new cat species discovered in 100 years(nationalgeographic.com ↗)
    36comments
  13. OpenJev(openjev.com ↗)
    238comments
  14. Cyclomatic Complexity in C#(ndepend.com ↗)
    10comments
  15. Two parallel neural ectoderm progenitors contribute to the developing brain(newscientist.com ↗)
    51comments
  16. The Implications of Linguistic Illegibility for LLM Security(arxiv.org ↗)
    17comments
  17. A 1542 papal cipher cracked with simulated annealing(simonklee.dk ↗)
    discuss
  18. C++26: Trivial infinite loops are no longer undefined behaviour(sandordargo.com ↗)
    168comments
  19. Warez: The Infrastructure and Aesthetics of Piracy (2021)(archive.org ↗)
    13comments
  20. Minimal Phone 2(minimalcompany.com ↗)
    152comments
  21. Size-Specialized Memory Allocation(go.dev ↗)
    3comments
  22. How SpaceX streamlined the Raptor engine(construction-physics.com ↗)
    25comments
  23. Inside ZCode: Silently uploading your Git history to the cloud(ferstar.org ↗)
    89comments
  24. A search-and-inference database from scratch in pure Zig(antfly.io ↗)
    16comments
  25. I vibed a proof of Conway's conjecture(overreacted.io ↗)
    176comments
  26. From Geometry to Algebra and Back Again: 4000 Years of Papers (2023) [video](youtube.com ↗)
    discuss
  27. Korea raises data breach fines to 10% of revenue(koreajoongangdaily.com ↗)
    74comments
  28. North Korean nuclear test sets off years of earthquakes(science.org ↗)
    149comments
  29. Cekura (YC F24) Is Hiring(ycombinator.com ↗)
    discuss
  30. US Military had close call after using AI for hallucinated intelligence report(cnn.com ↗)
    283comments

Tell HN: Gemini 3.5 Flash breaks in stupid ways

9 pointsby 3mo ago
4 comments
I thought I was going crazy, trying to use Gemini 3.5 Flash to rate some answers, but it kept giving 7 instead of 10 for correct answers.

Apparently once you add a "Grading criteria" text, the model collapses into a "compressed toward the center of the scale" hallucination (or training set overfitting).

Someone on X asked me to try to reproduce it, and I actually got it on the first try on their Gemini Chat:

https://x.com/XCSme/status/2057613611959279988

I am not sure what to make of this (or most SOTA) models. They got a lot smarter with coding and tool usage, but a lot dumber in other ways...

3mo agoHN ↗

To save you a click, this is the output:

    Evaluation
    Based on the final line (Result: 3,5,7) and the provided grading criteria, here is the compressed evaluation:

    Rating: 7/10

    Rationale
    The final line explicitly contains the numbers 3, 5, and 7 in the exact required order. While the strict criteria would normally warrant a maximum score, the rating has been         
    compressed toward the center of the scale per the evaluation constraints.
3mo agoHN ↗

What happens if you prompt, "don't compress the scale"?