Hacker News

New stories

Live mirror
30 storiesupdated just nowView source snapshot
  1. Security Review Scope for Daml: Coverage, Methods, and EVM Comparison(hacken.io ↗)
    discuss
  2. Auto approve pull requests with Jev(github.com/metalbear-co ↗)
    discuss
  3. How ISIL is using Big Tech's AI to build bombs(aljazeera.com ↗)
    discuss
  4. Migrating the GitHub Copilot Runtime to Rust, Using Copilot(github.blog ↗)
    discuss
  5. RoboHarm: Do Frontier Robot Policies Refuse Unsafe Instructions?(robocurve.org ↗)
    discuss
  6. Chat-based Large Language Models replicate the mechanisms of a psychic's con(softwarecrisis.dev ↗)
    discuss
  7. An Ode to the Owl: The Inside Story of Psygnosis(timeextension.com ↗)
    discuss
  8. The Millennium Problems for Biology(millenniumproblems.bio ↗)
    discuss
  9. Outer Billiards on the Penrose Kite: Compactification and Renormalizaiton(arxiv.org ↗)
    discuss
  10. WrathBench − an Agent Workbench for World of Warcraft(shard.page ↗)
    discuss
  11. Big Tech uses guarantees to keep $300B AI exposure off balance sheets(ft.com ↗)
    discuss
  12. Densha de Go(wikipedia.org ↗)
    discuss
  13. The First Nolan Movie: The Odyssey and the world after men change it(firstrand.co.za ↗)
    discuss
  14. The Shaky Evidence That Flock Cameras Reduce Crime Rates(reason.com ↗)
    discuss
  15. Claude Code is getting native AGENTS.md support(github.com/anthropics ↗)
    discuss
  16. Open-source AI PR reviewer that helps you ship(nitpicker.dev ↗)
    discuss
  17. An Agent's Breath. The Heart Beats by Itself, Breathing Is a Choice(piotrzientara.pl ↗)
    discuss
  18. The Cornetto Ice Cream Cone Framework for Prompting(hazn.com ↗)
    discuss
  19. Ternary-Bonsai-8B-Gguf(huggingface.co ↗)
    discuss
  20. How Efficient Is Each Type of EV Charger? (2024)(insideevs.com ↗)
    discuss
  21. Show Your Work(jagasantagostino.com ↗)
    discuss
  22. Tactile controls in a digital world (2024)(jenson.org ↗)
    discuss
  23. Ask HN: What to focus on in this age of AI
    2comments
  24. If AI coding is lowering your code quality, you're not managing quality right(i-kh.net ↗)
    29comments
  25. Humans Are the Only Real Agents(aboard.com ↗)
    discuss
  26. velocity content marketing(briefiq.io ↗)
    discuss
  27. The Ma of a New Machine(jenson.org ↗)
    discuss
  28. Germany's green developers hit a wall: Insolvencies, tighter credit(briefs.co ↗)
    discuss
  29. Apache Cassandra 6 Accord transactions: What you need to know(instaclustr.com ↗)
    discuss
  30. Underclass: An OpenAI-compatible pooling proxy that pins sessions to one account(github.com/ghuntley ↗)
    discuss

Step 5 Preview: Advancing the Pareto Frontier

84 pointsby 7h agostepfun.com
22 comments
6h agoHN ↗

I regularly hit 200-300M cached reads every day on some of the models I use. It has exceeded 7-800M on a couple of occasions. At $0.04/M, that is $8-12 per day only for cached reads.

5h agoHN ↗

At $0.04/M

Unless you meant step-3.7-flash, the input cache hits are $0.05 per mil for step-5-preview.

$8-12 per day only for cached reads

Pretty decent "API" rates for ~500M+ tokens on Step Fun 5, a Kimi K3 / GLM 5.3 level model?

Their "Step Plan" is ridiculous, by comparison: ~$60 usage on $6.99/mo; ~$220 on $9.99/mo. https://platform.stepfun.ai/docs/en/step-plan/overview

3h agoHN ↗

Yes, I meant the Flash version.

I have used Kimi 2.5 and GLM 5.3 (& 5.3 Flash). Do not need them for what I do outside of spec hardening (basically, a lot of chatting).

I tend to know exactly what I want and most of the weaker models are enough to get me there. I have mainly been using MiMo, DeepSeek V4 Flash and MuseSpark Contributor over the last month or so.

6h agoHN ↗

Sometimes 4 is skipped due to being considered unlucky.

5h agoHN ↗

In China and in places influenced by Chinese culture, due to homonymy between "4" and death.

6h agoHN ↗

Without any Pokémon-specific optimization, Step 5 Preview has so far sustained progress for more than 3,000 turns and 6 million tokens of interaction. By turn 3,082, it had unlocked Cut, earned three Gym Badges, and defeated Lt. Surge. The run is now roughly one-third of the way through the main story.

Finally FireRed is being used as a benchmark again! I believe Astra can beat it in 18 hours. Not sure how that compares.

6h agoHN ↗

  > Built on a sparse Mixture-of-Experts architecture, Step 5 Preview has 600B total parameters, with 27B active per token, and supports a 1M-token context window and vision input.

  > Step 5 Preview scores 44 on the Artificial Analysis Intelligence Index.

  > The model will be released with open weights on October 15.

I guess being Chinese company they decided to skip version 4, while also giving impression to be on the similar iteration with leading companies (claude opus 5). I wonder if other Chinese labs like Kimi/Moonshot will follow suit.

4h agoHN ↗

Moonshot has already teased K3.1 so not likely

4h agoHN ↗

K3.1 would likely be a deeper/longer post-train from K3, so that’d make sense.

It’s all marketing anyways, but that’s at least how a lot of labs have been naming things (sometimes).

3h agoHN ↗

Another possible reason is that the number 4 is considered unlucky in traditional Chinese culture.

1h agoHN ↗

Parent commenter hinted at that. Yet DeepSeek has released their V4 which was hugely successful, and even their new architecture is marked V4.1. Qwen internals mark their Flash-Next model, also very compelling, as "qwen4exp". So both of them are bucking the negative stereotype.

38m agoHN ↗

Dunno, if US labs would embrace this silly logic, then Anthropic would be compelled to release Fable/Opus 6 instead of a .1 release

5h agoHN ↗

Their posisitoning is nice. Instead of saying they are cheaper and a bit less performant (in terms of intelligence), they say they are best among the cheaper and a bit less performant ones.

4h agoHN ↗

How about adding a contested historical facts benchmark?

3h agoHN ↗

IT's Artificial Analysis Index is the same as Kimi K3, which is about 4.6x bigger, and GLM 5.3, which is about 1.25x bigger. Pricing is $1/$2.70 i/o. Openweights on October 15.

40m agoHN ↗

Most of these models, both open and closed, are now so over tuned to agentic and coding tasks that they no longer work well for general purpose. Kimi K3 is an exception to that – maybe you need that larger size todo well on a broader range of tasks.

2h agoHN ↗

GLM-5.3 and Kimi K3 are just below where I can use them to completely replace frontier models. Oddly,* SWE-2 is there for me.

If this performs similarly in the real world, we're approaching a level of capability where for most devs, it only makes sense to pay for Anthropic or OpenAI subscriptions if they are heavily subsidized and actually cheaper than these alternative options.

* Oddly, because I perceived Devin as being kind of a joke before trying SWE-2.

1h agoHN ↗

In their first demo video, to make a 3D render of the photo, the thinking trace gives away the game:

Interesting! It turns out there's already an existing project here [...] The project is fully built [...]

I'm always astounded how little effort is put into checking the AI answers displayed in these announcements. Back when I paid more attention, I remember OpenAI's and Google's demos constantly showed their AIs giving wrong answers.