Hacker News

Top stories

Live mirror
30 storiesupdated just nowView source snapshot
  1. How Hacker News ranking works: scoring, controversy, and penalties (2013)(righto.com ↗)
    28comments
  2. I built non-autoregressive decision models with RL a year ago(convaiinnovations.com ↗)
    240comments
  3. AI-generated posters don’t have to be horrible(john.hartnup.uk ↗)
    719comments
  4. Measure internet censorship. Contribute to the largest open dataset(ooni.org ↗)
    38comments
  5. Mayday Mysteries(maydaymystery.org ↗)
    3comments
  6. Compiler-style optimization for drawing via Skia(arxiv.org ↗)
    11comments
  7. Brood War Bench(swerdlow.dev ↗)
    62comments
  8. English: A vs. An(redblobgames.com ↗)
    99comments
  9. Two parallel neural ectoderm progenitors contribute to the developing brain(stanford.edu ↗)
    231comments
  10. Deodands put a price on objects that caused death(jstor.org ↗)
    11comments
  11. ZK-JPEG: Zero-Knowledge Image Editing and Compression(iacr.org ↗)
    6comments
  12. Show HN: CUA-S1 – A System One Model for Computer Use(github.com/trycua ↗)
    7comments
  13. UFO Series Home Page: "UFO" TV Series from 1970(ufoseries.com ↗)
    26comments
  14. You can defeat the Dream Devourer from Chrono Trigger using an int overflow(chrono.fandom.com ↗)
    8comments
  15. Android 17 is the first since 3.x to add new APIs without releasing to the AOSP(grapheneos.social ↗)
    645comments
  16. Suzanne Ciani's Buchla Cookbook(echo.orpheusinstituut.be ↗)
    20comments
  17. Tin: full-text search for Postgres(planetscale.com ↗)
    70comments
  18. The Lamentable Later Life of Lemmings(filfre.net ↗)
    discuss
  19. Supabase (YC S20) Is Hiring for OrioleDB(supabase.link ↗)
    discuss
  20. The Secret Life of Circuits(coredump.cx ↗)
    71comments
  21. New evidence for hidden chambers beyond Tutankhamun's tomb(nature.com ↗)
    36comments
  22. Black Holes or Black Hole Stars? Astronomers Spar over 'Little Red Dots'(quantamagazine.org ↗)
    44comments
  23. GPT-6 Astra Solves a WWI German Radio Cipher(prinzai.com ↗)
    166comments
  24. I think you should almost never use AI to write(erichgrunewald.substack.com ↗)
    106comments
  25. San Francisco Onion Futures Company(onionfutures.com ↗)
    152comments
  26. How to Write with an LLM(sockpuppet.org ↗)
    372comments
  27. Cloudflare Quick Tunnels(cloudflare.com ↗)
    308comments
  28. Microsoft director: AI scraping 'the largest theft of labor in human history'(tomshardware.com ↗)
    26comments
  29. Adventures in Microcontroller Circuit Debugging(bigmessowires.com ↗)
    8comments
  30. If math is more than proof, we need to better celebrate the rest of it(terrytao.wordpress.com ↗)
    231comments

Brood War Bench

109 pointsby 8h agobw.swerdlow.dev
62 comments
3h agoHN ↗

A ton of conversations about the game must be in the training set. I wonder, is there any way just from watching how they play, of telling if they tend to pick strategies that people complain or meme about online?

3h agoHN ↗

I wonder if there is a library to decipher brood war replay files. Perhaps an agent could learn by watching.

3h agoHN ↗

I haven’t checked for SC:BW but Blizzard has official parsers/replayers for SC2 on GitHub.

2h agoHN ↗

Not hard to build. I was shocked at how fast/easy this was to pull together.

2h agoHN ↗

At high levels of play Zerg is generally considered significantly stronger than the other races. So I generally will assume AI will tend to pick Zerg

1h agoHN ↗

Unless you Protoss and mind control the Zerg, and then are Zerg+Protoss.

24m agoHN ↗

This is not a viable strategy in competitive play. It's a huge resource/time investment with a dubious payoff.

1h agoHN ↗

I only played the StarCraft demo two decades ago. Why is that so? Rushing?

3h agoHN ↗

This is a great idea for a benchmark. Something all the benchmarks seem to be missing is strategy, tactical solutions in most of the benchmarks are all thats required but here requires actual long term thinking and tactical thinking, balancing and orchestration.

3h agoHN ↗

Any details about the harness the agents were given? I am curious what representation of the screen and world state was provided to the agents and what tools they had available.

2h agoHN ↗

Oh sorry I should be more clear on that. Will add to report.

For agent harness I did Claude Code, Codex, Grok Build. This was primarily a cost driven decision — I have a lot of free tokens and I didn't want to pay API prices for this.

For game harness I used minimal BW-API issue command and get observation apis as tools. I felt this was the most fair way to do it on my small scale.

In the future I would like to integrate code mode and multiple games/I think if it was a best of 5 where each agent could learn from its past games and build its own automations over time that would be much more interesting.

3h agoHN ↗

Unrelated to the benchmark...

I love StarCraft. I started playing it right from the beginning, most of my friends right now are from that era. I literally met people that have spread to almost every continent when I was in my early teens. We played at internet cafes and did not have access to the internet, that was priced differently...

I miss those days so much.

Everybody was from a different background back then, and nobody was anything other than a guy that plays StaCraft at the cybercafe... And now, we are in our 40's and I know Math teachers, history teachers, oil rig operators, software programmers, professional gamers, lawyers and more... hahah So crazy to think about it... and I know them, we talk, what a world.

2h agoHN ↗

Later we even listened to Eminem and played violent video games. Most of have never even been charged with a crime, much less abused someone.

2h agoHN ↗

I agree.

My first time playing StarCraft was at summer camp around a decade after it came out.

All the smartest people played it so I wanted to too. Great decision, I have been continually impressed with the people who StarCraft introduced me to.

2h agoHN ↗

Even at Burning Man, in the middle of the desert, there is a camp that hosts a StarCraft tournament every year (on the dustiest setups you've ever seen!) :)

1h agoHN ↗

That would be cool. It's Captain Pump's Raiders

2h agoHN ↗

My best greetings to all Starcraft Elders clanners of yesteryear :).

2h agoHN ↗

Oh boy, I've spent more hours playing it than I dare to admit. I won over 10000 battle.net games ... on just one of my several accounts :) When the SC2 beta came out, I played about 50-100 games and never bought the full game, because I knew it would be like heroin to me, and I was already an adult that had to take care of himself.

1h agoHN ↗

Same and Brood War is starting to have a bit of a resurgence. Not anything huge. When Battle.net servers are actually working, tend to play on the ladder a few times a week. Get crushed, but still one of the greatest competitive games ever made.

1h agoHN ↗

Yes! Being young is great, and it always seems like those days were the best.

1h agoHN ↗

A better benchmark might be asking the LLMs to write GO ai and then comparing that-- the issue is that there will be a HUGE difference in performance that depends purely on this game being in the LLM's training... but training a general LLM to directly play these games would be a waste of capacity and shouldn't be encouraged for benchmaxxing sake.

Programming an engine OTOH is a skill that is more general and they should all have.

Might be useful to have the target of the engine be some specific virtual machine that gets a strict cycle budget-- e.g. execution runs so many cycles, and result is read out of a specific memory address at the end (or when it terminates early).

1h agoHN ↗

Excuse me? Is this thing supposed to be a borderline SGI or not? We already know LLMs are good at spitting out code.

1h agoHN ↗

Playing these games autoregressively isn't even the right way to use the LLM for this task (unless it was trained to do so...). It's somewhat like having a creative writing bechmark but requiring that all the input/output be base64 encoded. It can do it-- but no guarantees on the results!

And it's also just bencmaxxing bait: you can get a huge improvement on the task by RLing on it, but make no improvement on anything else. Doing so would just waste model capacity.

If you could tell that every LLM was equally not being exposed to the task then you could justify it as a test of abstract reasoning, but you can't. So it ends up on how much go transcripts ended up in the training, which is ... not a very interesting metric.

2h agoHN ↗

Would be interesting if you could team a fast and slow agent together -- slow model can either act directly or maybe just communicate to the fast model.

2h agoHN ↗

I might open this up to a tournament if enough people want. Any interest?

2h agoHN ↗

I predict LLMs will reach superhuman level and beat even that model in the next 12 months

2h agoHN ↗

Starcraft is APM-dependent. Unless the latency will improve greatly in frontier reasoning LLMs (which is unlikely), it will remain a bit like knitting with an excavator.

2h agoHN ↗

I predict latency will improve greatly in the next 12 months to more than 4x speed on current frontier tasks

2h agoHN ↗

Yeah but Starcraft needs, like, 10-20x the APM these agents are doing.

2h agoHN ↗

I’m not convinced a lot of it can’t be solved with code mode.

Marine staggering for example seems like an ideal code mode task.

2h agoHN ↗

Yeah. Some of it may just be "thinking" less rather than faster token generation.

2h agoHN ↗

Should be possible to play this with Jev.

2h agoHN ↗

AlphaStar beat one retired professional by cheating.

AlphaStar won a showmatch against TLO. TLO was never one of the strongest players in the world. He had been retired for over three years by the time of the match. Google set the rule that their system would have human-like mechanics, but it played several times faster than any human, never issued a wasted action, had an inhumanly fast reaction time, issued commands with perfect accuracy using an API, and could see the entire map at once.

It was later released to the open ladder with more human-level mechanics. Even strong amateurs regularly beat it. I have beaten it myself. It was strong, but not even close to the level of the strongest human players. It had obvious and easily-exploitable deficiencies in strategy and building placement.

I think even the cheater version would have lost handily to Serral or any of the strongest players.

(It apparently beat MaNa as well as TLO, but those matches were never released to my knowledge. I see no reason to assume Google cheated less flagrantly in private than they did in public.)

2h agoHN ↗

Did it play by looking at screenshots and sending clicks, or was there other mediation/symbolization?

It sounds like it might have been actually played in real time, which would be very important to distinguish.

I have recently seen other harnesses letting agents play real-time games in what seems like discrete time slices, turning eg Portal into something turn-based https://www.youtube.com/watch?v=ruuGXFAmiOE

2h agoHN ↗

There's a bot data stream already; it's probably hooked up to that rather than screencap. Yes, I believe these were playing in real time.

30m agoHN ↗

It plays from a special API. Some things that are impossible with normal UI are possible with API, like selecting a unit under other units. Invisible units are also reported via the API.

2h agoHN ↗

There is currently a bot beating everyone on the ladder. Just watched it today on Artosiscasts yt channel.

1h agoHN ↗

If you read the TL thread on this, the bot is technically running at 240 APM, but its commands get split off into separate commands for each unit, so the APM looks higher.

I've watched several replays on this. It using map hacks is lame but it also doesn't always take advantage of them. It sometimes does respect its own fog of war. It mostly wins with incredible micro (kinda has to, it's macro kinda sucks).

Most of its few losses come from drops (it never makes turrets and doesn't know how to handle them), bad macro (blocking its own ramp) or just incredibly unconventional play from its opponent (which is certainly not in its training set).

31m agoHN ↗

Specifically, it can micro across multiple screens of engagement in a way that we can't as easily.

1h agoHN ↗

It is in fact APM limited, and its creator claims it was trained without the map hack so it should do just as well without it. I think the bot is VERY impressive and has already managed to beat several top Korean pros which is really godlike

1h agoHN ↗

I’m a little bummed it did have map hack on. Would love to see this play “legit”. It really reminds me of my pastime, chess. When you play bots, if you don’t mess up early, it can often feel like you’re winning in the early game even though you’re just hanging yourself slowly with the rope it’s giving you.

1h agoHN ↗

Bots have been unable to beat even pretty good amateur humans despite those advantages so it's still pretty interesting.

2h agoHN ↗

I think we're on the early days of games you connect with your agent to. Human + AI units one versus the other. Like knights with their horses. Not sure which is the horse..

2h agoHN ↗

I've been running these with friends recently, it's very fun. I ran irl bot tournaments for board games a couple times but the agent era opens up a huge amount of possibilities.

Must recently I built out a MMORPG puzzle box thing, I wrote a general game architecture doc but left the specific puzzle design up to Fable. Nobody is actively playing rn but I left it up at bot.willmorrison.net.

2h agoHN ↗

Back in 2010, during the early days of bwapi, there was a Brood War AI tournament held by the Expressive Intelligence Studio at UC Santa Cruz. It's interesting to see how different the approaches were back then, vs this or Deepmind's SC2 work.

https://web.archive.org/web/20091124210529/http://eis.ucsc.e...

There's a great contemporary Ars Technica piece by a competitor:

https://arstechnica.com/gaming/2011/01/skynet-meets-the-swar...

As an undergrad I did a project using genetic programming. It was not very successful, but it was a lot of fun.

https://tomisin.space/archive/starcraft-genetic-programming/

2h agoHN ↗

At last something that feels properly orthogonal to pelicans on bicycles.

2h agoHN ↗

Gemini wasn’t included, but I’m guessing its performance would have been similar to Grok’s performance despite having roots in DeepMind.

1h agoHN ↗

Wondering what it would look like if you allowed them to write scripts. You could throttle the number of clicks to make it interesting.

29m agoHN ↗

It can play doom in realtime so maybe it can play this.

24m agoHN ↗

I love this. Funnily enough StarCraft has influenced how I approach AI at a meta level

Protoss: powerful and expensive frontier coding agents you directly micromanage for the toughest tasks

Terran: versatile team comps of dedicated agent roles you can delegate well-defined tasks to

Zerg: massive swarms of specialist custom agents inside your apps that you evolve and optimise for speed and cost

Knowing every faction has its strengths and weaknesses helps me decide which tools to use for the job.