Hacker News

Top stories

Live mirror
30 storiesupdated just nowView source snapshot
  1. Breaking Up with Google Play: Why Conversations Is Now Free (gultsch.de)
    183comments
  2. PipePipe: NewPipe hard fork implementing SponsorBlock (github.com/infinityloop1308)
    42comments
  3. Fifteen years later, the Apple Cards origin story (lexontech.org)
    38comments
  4. Show HN: A Claude Code skill to analyze your chess games (github.com/brumar)
    16comments
  5. The Lost Atomic Update on Loongson CPU (jia.je)
    —discuss
  6. Modern Object Pascal Introduction for Programmers – Castle Game Engine (castle-engine.io)
    28comments
  7. I'm the Mom in That Viral Giants Clip. Let Me Tell You About My Husband (themomoftheyear.substack.com)
    7comments
  8. Revealing the details of how OpenAI agents hacked Hugging Face (swarmtraces.org)
    397comments
  9. Reflections on 1,000 Days of Math (gmays.com)
    7comments
  10. Automattic has a new board after failed attempt to put CEO on leave (techcrunch.com)
    35comments
  11. Analyzing Frontier Model Progress with My Favourite Game: Prince of Persia (blog.priyan.in)
    18comments
  12. We're gonna need a lot more mathematicians (terrytao.wordpress.com)
    367comments
  13. Make Claude your assistant in excalidraw (tangled.org)
    1comments
  14. Plan mode is dead (aymannadeem.com)
    426comments
  15. Floci: Locally emulating any cloud service (floci.io)
    22comments
  16. Banks and Credit Unions to Team Up Against Apple Pay Fees (macrumors.com)
    1comments
  17. Ollaya – Ollama for open-source, Jev-style decision models (ollaya.dev)
    133comments
  18. Is your Postgres migration safe or not safe? (safenotsafe.dev)
    31comments
  19. New Satellite Engine Could Use Earth's Atmosphere to Stay in Orbit Indefinitely (scitechdaily.com)
    2comments
  20. 16GB iPod Nano 3G Upgrade (tuckerosman.com)
    12comments
  21. Plunging test scores are a slow-moving catastrophe (economist.com)
    110comments
  22. Show HN: Jev Plays Pokémon Red (jev-pokemon.vercel.app)
    94comments
  23. Parsing Expression Grammar vs. Regexes: Building Org Parser in Lisp, Export HTML (jointhefreeworld.org)
    14comments
  24. A single function Jev-like wrapper for LLMs, including vision models (allanrbo.blogspot.com)
    36comments
  25. What even is an OS now? (sockpuppet.org)
    380comments
  26. Calculating atmospheric drag on satellites for a Cubesat [pdf] (osti.gov)
    7comments
  27. Ask HN: Who's still keeping a DOS machine up because the business depends on it?
    209comments
  28. Gravity seems holographic. What does that mean for reality? (quantamagazine.org)
    197comments
  29. The Murky History of Soviet-Born Tetris (mitpress.mit.edu)
    18comments
  30. Scientists build most accurate atomic clock (phys.org)
    23comments

Analyzing Frontier Model Progress with My Favourite Game: Prince of Persia

22 pointsby 20h agoblog.priyan.in
18 comments
1h agoHN ↗

The prompts provided are atrocious. It's amazing that the LLMs actually built something useful.

My first prompt was simple: "in this original code there 6502 assembly code for prince of persia, use the save level files and try to do it in c# console."

Do what in the what now in the console?

1h agoHN ↗

"now char coming but movement everything wrong ... prince is not in floor"

I'm assuming a language barrier on the part of the author. I wonder if the models would do better being prompted in the author's main language.

1h agoHN ↗

Or having another model proofread the prompts and write clearer instructions

1h agoHN ↗

Unsurprisingly it was buggy without the obligatory "Make no mistakes."

46m agoHN ↗

Yeah, I'm thinking the nigh-unreadable AI-speak we get these days makes a lot more sense if this is what they're training on. Or maybe the author has translated the prompts from another language?

38m agoHN ↗

It's so confusing how the actual article is in English, but the prompts are just gibberish. But LLMs are pretty good at deciphering gibberish; I often put our CEO's absolutely atrocious E-Mails into ChatGPT and tell it to explain wtf he wants from me.

Also, I feel like the LLMs would have done better if they had started from scratch each time, rather than being burdened by the output from the previous attempt.

1h agoHN ↗

A German magazine (c't) ran a Asteroids programming contest in 2008 - create a client with access to the emulator output of the game that provides keyboard inputs to control the game.

The highest scoring submission that won the contest had a high score of around 137k. Last week, I had GPT-6 Astra, Sol and Luna implement and hill-climb on this task, as I wanted to see how big the difference in smartness is. Luna implemented something, but never exceeded ca. 20k points, with a large variance. Sol got something in the area of the humans implementation.

Astra, which finished fastest, had a highscore of around 1.7Mm when the game seemed to fairly reliably crash. On the way, it disassembled parts of the ROM to extract information about the game.

I didn't do a ton work to document and measure the specifics, but it was very impressive.

1h agoHN ↗

If I remember correctly, back then all "good" entries in the contest solved the game by syncing with the RNG and could basically predict where and how new asteroids would spawn. The difference between entries was planning ahead movement and residence against control issues due to the contest setup (like inputs getting delayed). I wonder how Astra managed so much better, I thought the problem was already solved

38m agoHN ↗

Thanks! I went back, and had a look at both the contest rules and the highest scoring submission again. Turns out I benchmarked against infinite games, whereas the contest measured a 5min time limit. So, Astra solved a different problem.

Comparing strategies, though, Astra does something pretty different from the highest scoring submission. It doesn't predict the RNG, instead it reacts to the visuals on-screen, planning ahead by estimating velocity of all objects on screen. Unlike the winning solution, it actively flies the ship, whereas the winning solution basically only rotates and teleports. So Astra behaves more like a regular "perfect" player would, rather than one that breaks the PRNG.

1h agoHN ↗

Off tangent but the latest prince of Persia game is really fun

1h agoHN ↗

I played The Lost Crown and it's top tier, I assume you're talking about The Rogue?

58m agoHN ↗

It looks like one of those spam shooty game ads you get between Youtube Shorts.

Prince of Persia is a piece of contemplative, subtle, beautifully, artistically minimal motion puzzle art. God knows what that is.

37m agoHN ↗

It's one of the best recent metroidvanias. It deserves a lot more attention. I was surprised by how good it was.

1h agoHN ↗

I am going to try to build the original Prince of Persia using Swift as it is the best language for building a macOS game. Being a classic 2D cinematic platformer, Apple's native frameworks provide exactly what I need without the overhead of a massive cross-platform engine, and I love simplicity

So here goes my weekend, flag this and get a life

53m agoHN ↗

Share it when you're done! Even if it isn't complete complete.

1h agoHN ↗

Just one thing you could do is not drag Prince of Persia into this.

Please.