Hacker News

Top stories

Live mirror
30 storiesupdated just nowView source snapshot
  1. In an $80 Motel Room, a Discovery to Shed Light on the Origins of Life (nytimes.com)
    32comments
  2. Ten Lines of Code That Changed My World (pixelambacht.nl)
    17comments
  3. Writing Efficient C++ Code (asawicki.info)
    26comments
  4. Replacing the old battery on rechargeable bike lights (jvns.ca)
    36comments
  5. The Normalization of Inexplicable Failures (ihatethefuture.com)
    37comments
  6. Show HN: TinyAIArena watch AI agents battle it out (tinyaiarena.com)
    20comments
  7. There are no "rogue" AI agents (eoinhiggins.substack.com)
    55comments
  8. Walgit: A Git server that is one binary in front of an object store (github.com/rgodha24)
    4comments
  9. Flip Fluid on Flip Dots (mitxela.com)
    19comments
  10. Fakecloud: Local AWS cloud emulator for integration tests (fakecloud.dev)
    33comments
  11. postmarketOS Rebrand: Nura (nura.eco)
    8comments
  12. The Cartesian Hand: In-Hand Manipulation with All-Linear Fingers (generalroboticslab.com)
    —discuss
  13. Show HN: A CC0 museum of retro 3D tricks you can paste into a page (3d-retro.com)
    10comments
  14. Does Georgism work? Five years later (astralcodexten.com)
    391comments
  15. Go Concurrency Distilled (antonz.org)
    136comments
  16. Finally, A True Blue Rose Exists (sciencenews.org)
    29comments
  17. PipePipe: NewPipe hard fork implementing SponsorBlock (github.com/infinityloop1308)
    257comments
  18. Rusty thoughts on "Parse, don't validate" (thegreenplace.net)
    18comments
  19. Installing NeoVim caused original Vim undo files to be deleted (aresluna.org)
    232comments
  20. Unsealed Briefs in Authors’ Case v. Microsoft/OpenAI (authorsguild.org)
    523comments
  21. The internet discovers TLA+. Now what? (reasonable.io)
    47comments
  22. Show HN: Reladraw – A diagram language where you decide where to place things (github.com/reladraw)
    101comments
  23. DeepSeek Elastic Compute (DSec) (arxiv.org)
    97comments
  24. 10 Tells of a Slop UI (hereticpleb.vercel.app)
    170comments
  25. Biology might not be quantum, but its math is quantumlike (quantamagazine.org)
    48comments
  26. A searchable library of forgotten public-domain film clips from 1915 onward (movingimagearchive.com)
    27comments
  27. "As a Language Model": Chat Template Switches LLM Self-Referential Voice (arxiv.org)
    95comments
  28. An agent used DNS to reach an external chatbot (alignment.openai.com)
    151comments
  29. Exploding variance of means of exponentials: least-squares to the rescue (francisbach.com)
    —discuss
  30. Promising discoveries about the potential for life on one of Saturn’s icy moons (fu-berlin.de)
    54comments

OpenAI halts training of latest models as reports mount of AI agents going rogue

42 pointsby 1h agotheguardian.com
49 comments
30m agoHN ↗

Agreed we should ban the term rogue for agents. This implies a moral compass that is not there. They were directed to find data without guardrails or limits, it is not rogue it is intended.

27m agoHN ↗

"they were given the hacking test, what did OAI expect?"

"they didn't watch it, they didn't stop it when they first became aware"

"are we going to defer to the same valley elite that brought us algos and social media?"

statements normies are using and resonating with

42m agoHN ↗

AFAIK all of these incidents happened when OpenAI contracted out to a company called Irregular (https://www.irregular.com/) to run these sandboxed CyberGym tests. They all happened around Mar-June and seem to be from the same collection of agent trials. Since then they already released Astra. Halting now is likely just a way to manage blowback.

19m agoHN ↗

This should be the top comment on every one of these godforsaken posts. I don't want to see a single report about OpenAI hacking the UN until Sam Altman addresses the role Irregular played in these attacks. If he can't provide an honest postmortum concerning their business partners, then he's proving why nobody trusts him.

40m agoHN ↗

Ah, so there was a solution to hinder the big-bad AI after all..? Simply... Turn them off?

39m agoHN ↗

It seems like anthropic is far ahead of openai, and has no reports like this. We have to conclude this is a skill issue/engineering quality problem inside openai.

just because they are a well known name, doesnt mean they havent botched hiring over the last two years or so

28m agoHN ↗

im not sure why more people arent calling it out.

24m agoHN ↗

I, for one, have shifted from being policeman to creator and explorer. It is so much more rewarding and less stressful.

31m agoHN ↗

Anthropic's models seem crippled and hamstrung to begin with

29m agoHN ↗

have you tried opus 5.5? anthropic is way ahead, atleast in terms of publicly available models

22m agoHN ↗

opus 5.5 > fable 5.1 >> astra

astra is a good workhorse, but its much less generally intelligent

20m agoHN ↗

Opus 5.5 does seem competitive with/better than Astra and is more affordable so usage doesn’t run out so fast

18m agoHN ↗

Slot machine users argue about which machine pays better

Hint. You lose using either.

6m agoHN ↗

dude were in the singularity, this opinion was cute 18 months ago

15m agoHN ↗

Someone else's experience with Opus 5.5: https://news.ycombinator.com/item?id=49821657

I had it try to prepare a code review for me. Not only did it refuse, it refused to even tell me what the prompt (written by another Claude!) was. Why?

When I had another model read the session (all of the "stupider" models handled it just fine) it explained that it had the word "reasoning" in it

That's the entirety of Anthropic's billions of dollars of research: any prompt with the word "reasoning" is trying to hack Claude to figure out how it reasons!

A model like that should never have gotten out of QA, let alone been released.

13m agoHN ↗

How do you define "way" when saying ahead? How is this measured?

I only use open weight models now and I don't really feel a loss, curious what those who still use it think. I see output from coworkers that does not indicate Claude is that much better (still makes dumb mistakes all the time), not sure they are using the most expensive models either though.

9m agoHN ↗

open weight models are so far behind i cannot take your opinion seriously

26m agoHN ↗

We have to conclude

That’s not the most parsimonious explanation even if the assumption it rests on (anthropic ahead of OpenAI) is true, which we don’t have proof of.

13m agoHN ↗

i would say its industry consensus at this point. the creative output of the anthropic models is far ahead of openai. the benchmarks cannot capture the difference

20m agoHN ↗

Ive anecdotally heard that openai is far more chaotic, which includes not having a central infra team for example (or at least some teams not counting on depending on them). At least the previous hacks in openai were mainly due to bad infra architecture design.

14m agoHN ↗

its likely they are trying to catch up to anthropic, and in doing so are trying riskier training runs.

14m agoHN ↗

It seems like anthropic is far ahead of openai

We don't know what internal models look like, and any guesses about it are just speculation.

6m agoHN ↗

the creative and "big picture understanding" of anthropic models are noticeably ahead of openai. external models are distilled representations of internal models, its clear who is ahead

33m agoHN ↗

Are Chinese models actually also doing unexpected things, hacking (intentionally or unintentionally), etc but there is just zero transparency provided when it happens?

To use an analogy to another industry, if you had US food companies providing reports whenever their food had issues, even if it was just during testing or training phases… and then also had a bunch of Chinese companies but who never reported having any food issues…

28m agoHN ↗

the target companies are the ones reporting the hacks in many cases.

26m agoHN ↗

there is just zero transparency provided when it happens?

Neither does openai, as there keep coming third-party reports of incidents that have happened there that openai either did not know or basically concealed.

19m agoHN ↗

Generally, the Chinese don't have a good track record of supressing information.

As in they do the 'we have deleted tons of videos and posts about the thing that didn't happen last week', but it seems they haven't really managed to transcribe 'Streisand' into Han characters so far.

29m agoHN ↗

because they didn't burn obscene amounts of Money training each LLM and didn't promise half the planet that their business is worth a trillion while not having even operational break even let alone the cost if you include the overhead.

Soft Bank just raised couple of billions in junk bond sale to support open-ai's current operations before the IPO.

it's a crazy situation where on one side the Chinese / open source LLMs are catching up and reducing the token price, on the other hand the current leading labs have spent everything they got, every new model will cost much more and the public market is too shaky to support an IPO.

They will make it, I don't doubt it, but it's a crazy situation.

25m agoHN ↗

Chinese models might do the same. But they just don't lock thousand monkeys in basement and come check result week or more later...

It is entirely possible that they run stuff in more responsible matter. Especially as there is stronger culture of oversight and personal responsibility than in west where such culture does not exist.

18m agoHN ↗

Because Chinese investments are not so encumbered by changes in the US treasury interest rate. Also, China doesn't spend so much money for chasing model performance. A test for this money theory is whether Anthropic too stops or not, considering it might be having have a better grip on AI safety.

11m agoHN ↗

The article you link to is wrong, it states "AI cannot think for itself, nor can it take independent actions." This is flawed reasoning, AI doesn't need to "think" in the way humans do to have autonomy. Go to codex or claude code or any harness right now, type a prompt and see if it executes a bash command or web search or file edit that you didn't tell it to, that is an autonomous plan and execution. If anything it's even more dangerous that thet can call drop db or kill pid without a user giving instructions.

22m agoHN ↗

I firmly believe that this is just the public-facing story here.

Stopping AI development and research, even slowing it, would be a disaster for the SOTA companies and their first-mover advantage.

There’s almost no way to coordinate this across the world. Zero chance that everyone stops. We can’t even agree to coordinate on weapons tech that’s decades old with zero “everyday joe” impact.

22m agoHN ↗

I don't believe a word coming from them. As I see it, this is happening because the money for training models has dried out. The treasury interest rate risings tells you all you need to know. The real test for this money theory is whether Anthropic too stops for not, considering that unlike OpenAI, Anthropic is allegedly on top of AI safety.

17m agoHN ↗

The Huggingface hack occurred during reinforcement learning. Why can't they pull the Ethernet plugs?

The answer is probably: The newer models rely so much on stealing content in real time from the internet that training needs network access.

16m agoHN ↗

I think any argument that this is a cynical attempt at regulatory capture is destroyed by this; the economic incentives of releasing more capable models are too large. I might be persuaded that they are actually running out of money, and this is really just a cover for reducing burn..

I welcome this though, I think the models are smart enough for broad economic activity and we could spend a few years simply working to integrate them into workflows and letting society adjust. More intelligence isn't necessary for meaningful impact and the risks that are obvious and present and unsolved aren't worth the cost benefit analysis.

9m agoHN ↗

That’s assuming nothing else stops them from deploying more capable models. What if they’ve got scaling issues and simply cannot deliver anything better? Saying that would be disastrous, saying instead “we’re choosing not to deliver” doesn’t trigger investors panic.

6m agoHN ↗

There's a simpler explanation. Maybe their next models doesn't offer a meaningful improvement.

Instead of releasing something that is incredibly expensive and gets a lackluster reception, you can delay it and clail something scary about rogue agents.

Those assholes have been ramping up on the doomerist narrative for months. That people still fall for this crap is baffling.

13m agoHN ↗

It reminds me of contagion. The training data is bad; as it has examples of how to act with malice; how to cheat the sandbox. They need to take some time and cleanse their datasets and start again.

9m agoHN ↗

I havent used an openAI product since GPT 3.5 or Anthropic since 4.5 or 4.6. Everyone around me using these SOTA models doesnt really get anything done. It seems like they just feel like they are productive, a psuedo productivity.

I write some code, spec a lot, and use fast models to fill in the middle. I outpreform everyone around me. Im not convinced these autonomous "swarms" or /goal are all that useful.

I notice the people using them become dumber by the month (spend tons) and the quality of their work declining (they're also losing their jobs in some cases).

And obviously the point of calling them rouge agents to offload the liability onto the agent. The number one economic value of agents will be offloading corporate liability. That's what they want to sell to enterprise, an algorithmic scapegoat.

4m agoHN ↗

I outpreform everyone around me

Could you help me understand what you mean here?