Hacker News

New stories

Live mirror
30 storiesupdated just nowView source snapshot
  1. Teaching a World Model to Play Pokemon (nostalgia.dev)
    —discuss
  2. North Korea 'Likely' Behind $388M Hack of Crypto Exchange Bitget (gizmodo.com)
    —discuss
  3. Excel now supports multiple values in a single cell (techcommunity.microsoft.com)
    —discuss
  4. Analyzing Frontier Model Progress with My Favourite Game: Prince of Persia (blog.priyan.in)
    —discuss
  5. Learning the Bitter Lesson of Agent Harnesses [video] (youtube.com)
    1comments
  6. What one vLLM replica on an L40 can carry and how it fails (percentes.ai)
    —discuss
  7. The stupidity and arrogance of GNOME developers (felipec.wordpress.com)
    —discuss
  8. Finally, A True Blue Rose Exists (sciencenews.org)
    —discuss
  9. Using Claude Code: Spending your effort (twitter.com/trq212)
    —discuss
  10. Cringe and Corny (ambrook.com)
    —discuss
  11. Yegge's Paradox (willbeddow.substack.com)
    —discuss
  12. Show HN: Piloxa – an MCP server that sends USPS Certified Mail from your AI (piloxa.com)
    —discuss
  13. Tell HN: Chrome 154 corrupt Fontconfig cache and crash KDE Plasma
    —discuss
  14. Too AI; Didn't Read (tai-dr.com)
    1comments
  15. The Scarcity Premium: AI made images free; advertising as a solvency bond (danieldeboulay.com)
    —discuss
  16. Asteroids named after Tom Lehrer and 'Weird Al' Yankovic (slashdot.org)
    —discuss
  17. Show HN: I read ISO 273 off the standard because the online charts disagree (silviaai.dev)
    —discuss
  18. Anthropic's Claims over Its "Supply Chain Risk" Exclusion by Dow Rejected (reason.com)
    2comments
  19. Show HN: A free resume builder where the AI can rewrite but can't invent (payanai.com)
    —discuss
  20. More than a taken branch per cycle? (lemire.me)
    —discuss
  21. Of 1,364 four-year colleges, 17 admit fewer than 1 in 10 applicants (regrok.info)
    —discuss
  22. One Shotted (quarter--mile.com)
    —discuss
  23. Container Registry on the Edge (ewpratten.com)
    —discuss
  24. AI starter pack for the X-curious (xstarterpack.com)
    —discuss
  25. AI will grow beyond our control, we must instill good values before it does (kradle.ai)
    —discuss
  26. Making `wc` 20x faster with parallel state machines (aarol.dev)
    —discuss
  27. Writing Efficient C++ Code (asawicki.info)
    —discuss
  28. Time Measurement in Game Programming (asawicki.info)
    —discuss
  29. 2027 Tesla Semi First Drive (motortrend.com)
    —discuss
  30. Build Plugins for Claude (claude.com)
    —discuss

Yes, Claude can do nine loops

100 pointsby 2h agoanthropic.com
49 comments
2h agoHN ↗

I'm sure there's a simple answer to this, and I'm probably just missing it, but what happened with the supergravity one?

And how many runs did it take before this one? They say "in one shot" with nothing more than "keep going." But we only see the successful run, reported by the people who ran it.

2h agoHN ↗

Cool article.

While it’s possible that this is just a much more AI-friendly problem, I don’t think it’s just that: I think the technology has genuinely gotten better.

It has gotten better imo. The blog author mentions 2 cases at the beginning -- users who think that AI will be capped and those who think it will be uncapped. From my perspective, both are technically right -- AI is capped or technically has usually reached some sort of cap, until human innovation improves it. AI doesn't really improve itself on a grand scale so much as humans improve it.

In other words, AI can and does iteratively improve, but every single ceiling we've spotted and broken through so far came from human ingenuity or effort. It will likely continue to require it, regardless of how much it can do on it's own. In that regard, it seems as though all of this will inevitably be "uncapped, until it reaches a cap, and then likely it will eventually be uncapped by humans (again)". Because of this, AI will never perfectly fit neatly into an 'uncapped' or 'capped' bucket, as long as time continues moving and we continue solving issues as they crop up.

1h agoHN ↗

I feel like usage of AI alone should be enough to uncap it, so long as AI companies do the predictable thing and steal everything from everyone.

There are thousands of people figuring out how to make AI better for their particular thing. Way more than that walking the AI through taking things from problem to solution.

1h agoHN ↗

AI can totally identify optimizations within a paradigm without human intervention. What it isn't good at now is deciding to try totally new paradigms. Humans aren't great at this either, though it's more cultural in our case.

1h agoHN ↗

I'm an AI skeptic, or even anti-AI, and even I would have a hard time to admit it hasn't gotten any better.

In fact I will even admit it, yes AI has gotten better... but also how could it NOT? It has literally all the resources it can has, namely attention from everyone, everybody talks about it, a lot of workers, from state of the art researchers to annotators, all the material resources, from dedicated chips to networking to water to electricity, the entire available dataset of Human written and said thoughts, literally everything published and thus categorized.

AI literally has everything humanity can provide, it better be "better".

1h agoHN ↗

The blog author

I also liked the article and the author before this post. But let's be honest that this was paid work as part of Anthropic's PR campaign, not just a random blog post.

2h agoHN ↗

AI can now do really hard science homework all by itself without messing it up, which is kinda cool.

1h agoHN ↗

Yep. There goes "hallucinations are inevitable" crowd.

2h agoHN ↗

Just recently I used Fable 5.1 and Astra to solve and prove a long-standing mathematical problem I was always interested in: the shoreline search problem (a ship in total fog is at unknown distance from the shore [infinite line], what is the optimal trajectory?) A particular kind of logarithmic spiral was conjectured 33 years ago; a few days ago, I have obtained the proof. How routine it has become.

https://arxiv.org/abs/2609.24454

1h agoHN ↗

Cool, that was an interesting read on its own. I viewed the html version and skimmed it, I appreciated your illustrations.

1h agoHN ↗

What if the shore isn't a clean line?

2h agoHN ↗

Notwithstanding the duplication of these posts across social media, OpenAI, and Anthropic ("it's all the model, they just tell it to keep going" if you beat anything with RL enough ... it's going to do the thing)

Here's my issue with this post:

True to the spirit of the challenge, they didn’t use millions of dollars in computer power. They used Fable 5.1, working within Claude Science, a platform scientists can pay to use.

Okay, billions of dollars have been poured into these agentic LMs, right? Each training run to get the next increment is costing millions of dollars?

This feels like an obvious jab at Navier-Stokes, but where we get to shift the numbers around to hide where the compute actually is being spent ... compute is being spent. It's either being spent in amortization to make the search smarter ahead of time, during training, or its being spent after.

Also love: scientists get to pay Anthropic to work within their special science harness to do science. That's exactly what I dreamed of doing when I pursued physics in undergrad, one or two companies holding the keys to "progress" for a monthly subscription price.

1h agoHN ↗

The training costs get rolled into the usage prices. It's not hiding anything to talk about the cost of usage without the cost of training, any more than I'd be hiding something by talking about the price of a $1,000 CPU without mentioning the tens of billions of dollars in R&D and infrastructure needed to create it.

1h agoHN ↗

Also love: scientists get to pay Anthropic to work within their special science harness to do science. That's exactly what I dreamed of doing when I pursued physics in undergrad, one or two companies holding the keys to "progress" for a monthly subscription price.

I get this anxiety, and am largely an AI skeptic, but at one point there were only a handful of computers in the world too (same for batteries, or engines, or crucibles, or stills -- it goes way back), and the organizations that had them had a stranglehold on progress in the field, as did the small number of companies who knew how to make them. It got better as they got cheaper and more plentiful.

I guess my point is that there are more important anxieties to feed when it comes to LLMs and the current state of the world.

1h agoHN ↗

I get this anxiety, and am largely an AI skeptic, but at one point there were only a handful of computers in the world too

I don't see how that makes the state of affairs any better.

1h agoHN ↗

This feels like an obvious jab at Navier-Stokes, but where we get to shift the numbers around to hide where the compute actually is being spent ... compute is being spent. It's either being spent in amortization to make the search smarter ahead of time, during training, or its being spent after.

I think that argument is recursive? These posts aren't very complicated for either of us, but they're written on devices that are fabricated with billions of dollars of semiconductor equipment. At what point do we just acknowledge that we stand on the shoulders of giants?

To me, the distinguishing factor is that the expense not special-purpose but upfront. The model here is trained without foreknowledge of what problems it will solve. Solutions like nine loops are genuine expressions of a pre-existing model capability, even if that capability has not pre-existed for very long.

1h agoHN ↗

Oh I agree with you! I just don't see "we're standing on the shoulders of giants" in most of these marketing blog posts?

If I'm wrong here, I'd love reference links. I think of these companies as trying to inspire the idea that Claude (or GPT) are these special alien entities, in a sense?

1h agoHN ↗

What a weird thing to be butthurt about. You don't have to pay them. Do it the old fashioned way, with elbow grease and pots of coffee. Or invest in local AI.

Or embrace the future and realize that you couldn't imagine everything that would unfold, when you pursued your undergrad.

How you gonna get mad about all this?

1h agoHN ↗

There was a time when a generation of hackers got (rightfully) worked up about the Microsoft tax for every PC sold. It's not hard to see that somebody gets mad if they believe that an Anthropic/OpenAI tax is about to become the standard when doing any serious Physics/Math/... research or writing code.

1h agoHN ↗

Divide the compute of your argument (pre-training, training, post-training) through everything this model now can do and it will not look that bad at all.

1h agoHN ↗

Also love: scientists get to pay Anthropic to work within their special science harness to do science. That's exactly what I dreamed of doing when I pursued physics in undergrad, one or two companies holding the keys to "progress" for a monthly subscription price.

The moat is very limited. Harnesses aren't crazy hard to engineer. Open models are quite capable.

1h agoHN ↗

Agree with you, but not sure what form the arms race will take as the months go by (between harness engineering, RLing models within harnesses, etc)

2h agoHN ↗

I’ve heard from smart, well-informed people who are confident that AI is a few years away from superintelligence, and that superintelligence will be capable of truly terrifying things. And I’ve heard from smart, well-informed people who are equally confident that LLM-based AI is close to a ceiling, that models like Claude won’t even be able to do impressive work in physics, let alone conquer the world.

This is framed as an "or", as if they're contradictory.

IMO, Both of these statements are true.

1h agoHN ↗

i.e., we're entering RSI and LLMs will help to build their more capable successors under a new paradigm? Fair point, but most people saying #2 surely mean to say LLM development has been a dead end and hasn't gotten us substantially closer to AGI

1h agoHN ↗

normally i'd agree, for "most people"

but the author explicitly distinguished between "AI" and "LLM-based AI", and they work for an AI company, where making these distinctions are really important

1h agoHN ↗

  > As it turned out, the result wasn’t all that far away for humans either. A few days after I heard from Anthropic, we heard from Song He, an amplitudeologist at the Chinese Academy of Sciences in Beijing. Song’s group had already gotten the majority of the result. They’d used some AI assistance, based on GPT-6, but not the kind of one-shot almost human-less approach Anthropic used.

Please have AI come up with something no human is also about to solve?

This gets me wondering why ai labs aren't proposing their own millenium prize type challenges.

1h agoHN ↗

Please have AI come up with something no human is also about to solve?

No human-only was close to Navier-Stokes. The team that was close was also using AI.

1h agoHN ↗

Fair point, and i think that's fine. AI labs could just not jump themselves into such problems and let humans have a go at them, whether ai-assisted or not. As in this case with a regular budget one researcher could've access to.

1h agoHN ↗

What is it about telling Claude "I'm only a biological human and now I have to go to sleep" that will get it to keep cranking away for hours. I have good success with similar statements and usually come back off several hours to find that Claude has managed to babysit several computational runs.

It's interesting to note that the expensive part of this experiment was the Claude operations expense. For me I find that Claude is a small fraction of my cost with most of the bill attributed to computers to run simulations instead of the AI to monitor and tweak the simulations.

1h agoHN ↗

I think it's mostly just what it is on it's face; You're telling it to be independent for a while. The default prompts start in an interactive mode where it checks in and verifies requirements a lot.

1h agoHN ↗

How are you guys able to get any of these agentic loops to actually run to completion before you run out of usage for the week?

1h agoHN ↗

Smart model in charge, planning & orchestrating the dumb models doing the work.

Often dumb models from other plans & subscriptions. I like to choose whatever model is on sale on opencode-go.

YMMV drastically based on what the actual work is. Dumb models can't handle everything. And of course I have no idea how much of a backlog of work you're feeding it or anything. Some people have enough work to exceed any plan, and do the work inefficiently to boot.

Also, you can just buy as much usage as you want at API rates.

1h agoHN ↗

First of all, there is definitely value addition with the LLMs in almost every field in some ways.

What bothers me the the marketing angle which invites skepticism and criticism

Anthropic invited Matt von Hippel to write this post and compensated him for his time.

If you are paying some one, tailoring the discussions then the end result is always going to be biased one showing yourself as the winner. I understand its somewhat organic, still the ratio of marketing and science needs to be balanced. Marketing has to be correct and the results/outcomes should be reproducible

1h agoHN ↗

only scratching the surface. The impact of AI might be significant/belligerent/virulent

1h agoHN ↗

I don’t understand what the model actually did here or how it was verified, and I feel like that’s a general problem with all these “AI did breakthrough science” posts. I understand that I’m not an expert in these fields so there’s only so much I can pick up from a blog post about it, but these posts are basically always just “trust us, it did the thing.” I mean surely you could easily share the conversation trace and any artifacts/code it produced so peers in the field can verify the work.

These scattering amplitude formulas are hard to compute, so hard that physicists almost always use approximations. They do partial calculations, cut off at a specific number of “loops,” a measure of how complicated interactions between particles are allowed to get. The more “loops” they include in their calculations, the closer they get to the real answer, and the harder, computationally, the calculation is to do.

I don’t even understand what type of solution we’re describing here, is it a formula? A program? A Lean proof?

1h agoHN ↗

Would you understand it if it had been done the old-fashioned way? I think the problem is the part where you (and I) are not an expert in the field. It's extremely rare for the solution to an unsolved problem like this to be understandable to someone outside the field.

1h agoHN ↗

It may have gotten a boost from using Python, and not Maple (Lance’s favorite program for math) or Mathematica (mine), and it may have used much better software engineering practices than we would have, but not super-intelligently so.

This is the crux of the matter here. Physicists are physicists. They are not software engineers. I read physics between 2001 and 2005, and the programming language they had us use for all of our assignments was FORTRAN 77. They were still trying to decide if it was ok to move students on to FORTRAN 95.

FORTRAN is pretty performant - if you know what you’re doing. Very, very few people did - and they, me, went on to have careers in software, not physics. I remember some FITS (astronomical image format) processing software someone was using to calculate ephemera - and it took DAYS to run over a few thousand images and produce an output. I sat down with it, screwed around for an afternoon, and it produced a result in under a minute - so much faster that I honestly thought I had broken it - but I hadn’t. It was just terribly written by someone who was excellent in their domain and terrible at writing code.

I see there being an enormous opportunity here, in AI providing scientists with software that isn’t diabolical.

1h agoHN ↗

Exactly the same case here. A few years ago, pre-ChatGPT, my wife was working with a massive database in SPSS or R (can't remember which). Now these "databases" are really just huge CSV/TSV files, with millions of records, one record per line. She was applying some very complicated mathematical functions to every single row, and she was frustrated because it was going to take days to apply a particular calculation she was working on to the whole database.

So I took a look for about 30 mins figuring out the code (I don't know SPSS or R), and it turns out she was doing some calculations that were identical for every single line, regardless of that line's fields. So I pre-calculated a few things, prior to starting the per-record-loop, and then just used the pre-calculated values every time. After that, every run took less than 5 minutes. She's amazing at math and economics, but optimizing low-level code is not something she has done in the past.

1h agoHN ↗

Yeah, very similar. For each and every image, the code loaded every image into memory, several times each, then did a pixel-by-pixel walk through the entire image, identified all objects, and then compared all of the objects to all of the objects from all of the previous images (which it would not use the calculated result for, but rather recompute that too) to establish which one was moving - or which one was staying still.

The main fix was not making it hideously IO bound, as not only was it loading every image multiple times, it was then grinding away in swap as we’re talking tens of gigabytes of raw image data, in an era when 1gb of ram was a lot - and then making it spiral out from the last known location of the object of interest rather than brute-forcing it. My solution wasn’t even optimal, as it was literally just an afternoon of tooling around as a favour.

1h agoHN ↗

Honey, it's time for your daily Anthropic PR piece!

1h agoHN ↗

I want to see AI beat a 4x strategy video game or a roguelike. Let me see a 20+ win streak on the hardest difficulty in Balatro or Slay the Spire. Let me see AI beat Civilization against the best human players.

Every time I mention this someone assures me that it's possible, and they point to simple board games that computers excel at, or real time strategy games where proper use of APM and clicking accurately go a long way. But I haven't seen AI succeed at any decision-focused game where describing the rules requires more than one minute.

I want to see what AI can do in, not simple, and not complex, but complicated toy environments, where decisions are all that matter.

11m agoHN ↗

This counts. Thanks!

I'm a bit skeptical of their "2 seeds in a row!" boast. Last time I investigated a claim like that I found the seeds were cherry picked. This was way back in the OpenAI Gym days though (remember when OpenAI was open and just doing goofy research like OpenAI Gym?), their leader boards had some amazing claims about certain RL solutions, but when I ran them myself on new seeds they were far worse than claimed.

1h agoHN ↗

You can get an OpenAI subscription and have it play games right now, today.

A lot of influencers are trying it. You can set it loose on games like Slay the Spire without any training and it can win: https://www.youtube.com/watch?v=9bDG0uuHM2w

That's a general purpose LLM. Actually training a model for a game like that would be old news.

I do have to laugh at how the goalposts keep moving to higher and higher levels like "I need to see it win 20X in a row at the highest difficulty! Why has nobody shown this?"

16m agoHN ↗

Humans can do that, so that was my goal post. I don't think I've changed it since this is the first time I've proposed a goal post.

Thanks for the video link. I tried to find one like this recently, but the video you linked wasn't in the results I looked at.

1h agoHN ↗

Show that an AI can take the kinds of computer resources an academic has access

I don't know how this can be verified, but I'm not convinced by the evidence in the post. It sounds like the final solution can be ran in a week on 96 cores. But how much compute was used attempting the problem? How much inference? 'Trust me bro?'

I think these solutions are worth millions to these labs, and I think that's what they're spending on them. Far more resources than have been directed at mathematicians and physicists to solve the same problems.

1h agoHN ↗

Even when a goal is simple and well-defined, sometimes it’s going to look much less achievable to experts than it actually is

I think it’s a compelling idea, that even if AI never achieves superintelligence, there might still be enormous value in their ability to relentlessly pursue a solution for much “longer” (relatively speaking) than a human. Even if they never let us pick the highest hanging fruit, if they instead let us pick every single low hanging fruit, that’s still a massive win.

46m agoHN ↗

in terms of UX for blogs, I don't understand why the summary was not posed and then had a click through to the full article.

I really wish this would become the default, it would potentially remove ad-rev but it is so much easier to not only get people knowledgable about the content but intrested in the subject matter to dig deeper.

re: https://imgur.com/a/Ko6rAbO

Note: cant post ai text or comments get auto flagged