Hacker News

Top stories

Live mirror
30 storiesupdated just nowView source snapshot
  1. Dream-RSI: Recursive Self-Improvement through Evolving Worlds(arxiv.org ↗)
  2. Mistral X Mozilla: Private, Multilingual AI Browsing(mistral.ai ↗)
  3. Small Programming Tricks(will-keleher.com ↗)
  4. Introducing System One Models and Jev(typesafe.ai ↗)
  5. Show HN: An e-ink frame that hears birds and draws them as 1800s illustrations(github.com/arnegiacomo ↗)
  6. Tell the speakers that you liked their talks(ohhelloana.blog ↗)
  7. Hackers Got Inside a Flock Camera. Its Data Shows How the System Works(wired.com ↗)
  8. How Big Are Factorials?(thegreenplace.net ↗)
  9. Can we stop with the uptime percentages?(jim-nielsen.com ↗)
  10. Apple Reference Image: A New Approach for Verified Photography(security.apple.com ↗)
  11. The Google Play app review process now regularly takes longer than a week(gultsch.social ↗)
  12. Scaling Golang CI by Replacing actions/setup-go(cloudx.ai ↗)
  13. Douglas Adams and the exterminated Doctor Who adventure(bbc.co.uk ↗)
  14. Kyber (YC W23) Is Hiring a Forward Deployed Engineer(ycombinator.com ↗)
  15. Original Sony PlayStation 2 security chip 'broken wide open' after 26 years(tomshardware.com ↗)
  16. Anatomy of a Texture(agentlien.github.io ↗)
  17. An update on Wayback Machine access(blog.archive.org ↗)
  18. Salesforce Global Outage(salesforce.com ↗)
  19. Show HN: How Stale Is Your AI? Release age and training cutoff for 20 models(stale.jock.pl ↗)
  20. Measuring Gauss-Seidel loop-carried dependency and fixing it via loop unrolling(loiseaujc.github.io ↗)
  21. Show HN: I made a flight simulator, except you're just a passenger(inflightsimulator.com ↗)
  22. Doing Everyone Else's Job(yosefk.com ↗)
  23. Gemini 3.8 Live and 3.8 Live Extended Thinking(blog.google ↗)
  24. PS5 Linux lead quits: "a bunch of noobs using LLMs" that "they don't understand"(frvr.com ↗)
  25. Why I'm still bearish on LLMs after Navier-Stokes(dank.systems ↗)
  26. DeepSeek v4.1 Flash Is Now Our Best Hacking Model(enclave.ai ↗)
  27. OpenAI expands ChatGPT ads with Sponsored Agents(openai.com ↗)
  28. Intelligence per Watt: Measuring Intelligence Efficiency of Local AI(arxiv.org ↗)
  29. German Rheinmetall open-sources its Battlesuite connected weapon system protcol(rheinmetall.github.io ↗)
  30. A software thing I built: GPS on a 25MHz 486-SX(vcfed.org ↗)

DeepSeek v4.1 Flash Is Now Our Best Hacking Model

88 pointsby 3h agoenclave.ai
24 comments
1h agoHN ↗

Apparently they updated it perhaps based on your comment?

. The accepted runs cost $4.65. Failed attempts and replacement runs increased the complete cost to $5.14.

56m agoHN ↗

Oh - I see that now. It's possible I missed it originally - I did read the article but I was skimming quickly. Mea culpa, if so.

1h agoHN ↗

You can have 100 runs for the price of the Claude Max plan?

1h agoHN ↗

I find this - or perhaps the title - a bit surprising.

I've benchmarked GLM 5.3 and DSv4.1-F on my fully-annotated decomp of the Nintendo 3DS's kernel, which I have a good mental understanding of, tasking them to find vulns and other bugs (in Max mode w/ subagents). GLM 5.3 founds almost all the vulns in 30min for $22, while DS only found one vuln for $2 in 40min.

Perhaps DS works better where targets have low-hanging fruits than can be found fast?

54m agoHN ↗

In your example, couldn't you parallelize DS's work more ? You could have 11 times as many agents for the same price.

35m agoHN ↗

DS has perf issues if you parallelize it heavily (24+), especially when the context window is above the limit, on a single machine (with custom llm gateway): it fails 4x more often, and is 2x slower than gpt-5.6.

14m agoHN ↗

Both DS and GLM had the same numbers of subagents, 5 or so.

But, well, number of subagents doesn't make a difference if model is dumb (GPT 5.4 High, in May,in Chat mode outperformed what I see with DS4.1-F).

That being said, pricing model makes a huge difference for "find at least one" tasks: with API/PAYG if you have a chance to save 90%, you go for it, whereas with subscriptions it is optimal to burn all your remaining allowance right before reset

39m agoHN ↗

Makes sense, GLM is a lot better than DS v4.1, I found the same results in other domains.

Given how fast and cheap DS is, it's just an ideal model with enough "IQ" to let it loose. Another thing they left out of the article, DS becomes really good with if provide custom tools for the task, on it's own it's mediocre.

17m agoHN ↗

How does GLM 5.3 Flash rank against it's big brother and the latest Deepseek?

38m agoHN ↗

I think it really depends on what the data DS was fine tuned on. If your use case is very specific, it wouldn’t have distilled that knowledge well.

30m agoHN ↗

What I find surprising is that DS Flash can do it at all.

I love DS flash, it is an amazing workhorse to implement plans created by more robust models (such as GLM). But a more fair comparison would be of DS Flash with GLM Flash.

29m agoHN ↗

Thing is that GLM 5.3 is many multiples the cost to run, and slower.

I have good results with DS4.1 flash because I can iterate faster. I either provide it with correction, or it discovers its failures via the harness. And seems to respond well to empirical evidence rather than go in circles.

So it might need some prodding, but it's likely in this case it was able to brute force after several runs and collecting some evidence.

18m agoHN ↗

A little off-topic: where does one use those models such as GLM or DS for this kind of reverse engineering tasks? I think I read many of them refuse to help with tasks like those on their official platforms.

3m agoHN ↗

GLM 5.3 doesn't seem to refuse vuln research (which it classifies as "audit") and is good at it.

Therefore you use for offensive cybersecurity tasks because Daybreak Red/Mythos is pure unobtainium for us mere plebians.

DB Blue thankfully exists, but I suspect you risk a ban if you use it with codebases you neither own nor use

5m agoHN ↗

If you are talking about publicly known vulns, it's a bit moot since they should be in the training sets. If not, you just burned the vulns to that inference provider's training data (and any intermediary), and future benchmarks will be meaningless.

47m agoHN ↗

When the history books are written and all is said and done, the hubris of this moment where all the American labs decided to punk their investors and join hand in hand in agreeing to let the Chinese win forever is going to be the main story.

42m agoHN ↗

Win what? The race to the bottom always has this competitive language.

“If we ban CFCs now the Chinese will win!”

“If we ban chemical weapons, nuclear weapons, etc etc our enemies will triumph! They won’t stop!”

“If we switch to biodegradeable plastic then our rivals will have an advantage.”

“If we dont externalize the costs to our population, then they will, and then will win!”

I think workflows can do the job agents do, 20x cheaper and more predictably and safely. They can completely displace agents, just as HFCs displaced CFCs and then we were able to ban CFCs and phase them out through international COOPERATION. The language of COOPERATION is what saves us vs COMPETITION is all about cutting corners and externalizing costs. Google the Montreal Protocol, Geneva Conventions, Nuclear Non Proliferation Treaty, Unleaded Gasoline etc etc.

Agents have got to be marginalized. They are just popular because the labs need to make a ton of money for their investors and recoup their massive spending on training models.

41m agoHN ↗

Those comparisons are pretty irrelevant

You can't compare banning football to banning genetic experiments and say "they are both bans and therefore directly comparable"

13m agoHN ↗

When the history books are written and all is said and done, the hubris of this moment where all the American labs decided to punk their investors and join hand in hand in agreeing to let the Chinese win forever is going to be the main story.

Come on. Lecturing about hubris when your message is damn the consequences, full speed ahead?

If it's a race to build the torment nexus, or a race with a nonzero chance of building the torment nexus by accident, I don't care about winning.

44m agoHN ↗

Seems pretty bold to claim deepseek is the "best hacking model" while providing zero comparisons to other models...

33m agoHN ↗

Notice the qualifier "our", that is the one they have access to.

25m agoHN ↗

What if there would be a separate category for distilled models?

27m agoHN ↗

DeepSeek is underrated. Basically all Chinese models are good enough for day to day coding at this point.

The 2 trillion dollar ROI on anthropic alone?

Good luck with that.