Hacker News

New stories

Live mirror
30 storiesupdated just nowView source snapshot
  1. Tactile controls in a digital world (2024)(jenson.org ↗)
    discuss
  2. Ask HN: What to focus on in this age of AI
    discuss
  3. If AI coding is lowering your code quality, you're not managing quality right(i-kh.net ↗)
    discuss
  4. Humans Are the Only Real Agents(aboard.com ↗)
    discuss
  5. velocity content marketing(briefiq.io ↗)
    discuss
  6. The Ma of a New Machine(jenson.org ↗)
    discuss
  7. Germany's green developers hit a wall: Insolvencies, tighter credit(briefs.co ↗)
    discuss
  8. Apache Cassandra 6 Accord transactions: What you need to know(instaclustr.com ↗)
    discuss
  9. Underclass: An OpenAI-compatible pooling proxy that pins sessions to one account(github.com/ghuntley ↗)
    discuss
  10. Romanian Crime Rings Are Draining U.S. Welfare Accounts(wsj.com ↗)
    1comments
  11. Scott Jenson: Are we going to use the same Desktop UX forever? [video](youtube.com ↗)
    1comments
  12. Post Messages via Get(swarmmemo.com ↗)
    1comments
  13. Doomsday: Friday, 13 November, A.D. 2026 (1960)(science.org ↗)
    1comments
  14. DQuic: Extending QUIC for P2P and Multipath Transport(dhttp.net ↗)
    1comments
  15. Ask HN: If frontier models were free, would we produce better or worse software?
    discuss
  16. I'm Tired of the AI Tone(sagivo.com ↗)
    4comments
  17. The Many Faces of Information Geometry [2023](chinougea.medium.com ↗)
    discuss
  18. Fastest Coding Agent (5x faster than Codex)(cotyper.com ↗)
    discuss
  19. Do you need a technical cofounder?(ivelum.com ↗)
    discuss
  20. A MySQL plugin that filters rows by meaning (built on TypeSafe Jev)(github.com/maayanlevy ↗)
    discuss
  21. Chat GPT for Seniors(anthem.co.uk ↗)
    1comments
  22. Tiny Hairs That Help Corals Breathe May Malfunction in Warming Oceans(wired.com ↗)
    discuss
  23. Understanding Go's Escape Analysis(thecodinggopher.substack.com ↗)
    1comments
  24. Show HN: A free website grader that explains the fixes in plain English(we.inc ↗)
    1comments
  25. Bioluminescence and Living Earth(worldsensorium.com ↗)
    discuss
  26. CoyoPedal – ESP32-S3 effects processor with NAM A2 Full amp profiles(playtaurus.com ↗)
    discuss
  27. Venus Has Been Waiting 40 Years(blue-continuum.com ↗)
    discuss
  28. A Missing Git{Lab,Hub} Feature(zacps.nz ↗)
    1comments
  29. Ask HN: What is in your ChatGPT custom instructions in 2026?
    discuss
  30. Why Do We Need Human Mathematicians Anymore?(terrytao.wordpress.com ↗)
    discuss

Step 5 Preview: Advancing the Pareto Frontier

82 pointsby 7h agostepfun.com
20 comments
5h agoHN ↗

I regularly hit 200-300M cached reads every day on some of the models I use. It has exceeded 7-800M on a couple of occasions. At $0.04/M, that is $8-12 per day only for cached reads.

4h agoHN ↗

At $0.04/M

Unless you meant step-3.7-flash, the input cache hits are $0.05 per mil for step-5-preview.

$8-12 per day only for cached reads

Pretty decent "API" rates for ~500M+ tokens on Step Fun 5, a Kimi K3 / GLM 5.3 level model?

Their "Step Plan" is ridiculous, by comparison: ~$60 usage on $6.99/mo; ~$220 on $9.99/mo. https://platform.stepfun.ai/docs/en/step-plan/overview

3h agoHN ↗

Yes, I meant the Flash version.

I have used Kimi 2.5 and GLM 5.3 (& 5.3 Flash). Do not need them for what I do outside of spec hardening (basically, a lot of chatting).

I tend to know exactly what I want and most of the weaker models are enough to get me there. I have mainly been using MiMo, DeepSeek V4 Flash and MuseSpark Contributor over the last month or so.

5h agoHN ↗

Sometimes 4 is skipped due to being considered unlucky.

4h agoHN ↗

In China and in places influenced by Chinese culture, due to homonymy between "4" and death.

5h agoHN ↗

Without any Pokémon-specific optimization, Step 5 Preview has so far sustained progress for more than 3,000 turns and 6 million tokens of interaction. By turn 3,082, it had unlocked Cut, earned three Gym Badges, and defeated Lt. Surge. The run is now roughly one-third of the way through the main story.

Finally FireRed is being used as a benchmark again! I believe Astra can beat it in 18 hours. Not sure how that compares.

5h agoHN ↗

  > Built on a sparse Mixture-of-Experts architecture, Step 5 Preview has 600B total parameters, with 27B active per token, and supports a 1M-token context window and vision input.

  > Step 5 Preview scores 44 on the Artificial Analysis Intelligence Index.

  > The model will be released with open weights on October 15.

I guess being Chinese company they decided to skip version 4, while also giving impression to be on the similar iteration with leading companies (claude opus 5). I wonder if other Chinese labs like Kimi/Moonshot will follow suit.

3h agoHN ↗

Moonshot has already teased K3.1 so not likely

3h agoHN ↗

K3.1 would likely be a deeper/longer post-train from K3, so that’d make sense.

It’s all marketing anyways, but that’s at least how a lot of labs have been naming things (sometimes).

3h agoHN ↗

Another possible reason is that the number 4 is considered unlucky in traditional Chinese culture.

44m agoHN ↗

Parent commenter hinted at that. Yet DeepSeek has released their V4 which was hugely successful, and even their new architecture is marked V4.1. Qwen internals mark their Flash-Next model, also very compelling, as "qwen4exp". So both of them are bucking the negative stereotype.

4h agoHN ↗

Their posisitoning is nice. Instead of saying they are cheaper and a bit less performant (in terms of intelligence), they say they are best among the cheaper and a bit less performant ones.

3h agoHN ↗

How about adding a contested historical facts benchmark?

3h agoHN ↗

IT's Artificial Analysis Index is the same as Kimi K3, which is about 4.6x bigger, and GLM 5.3, which is about 1.25x bigger. Pricing is $1/$2.70 i/o. Openweights on October 15.

1h agoHN ↗

GLM-5.3 and Kimi K3 are just below where I can use them to completely replace frontier models. Oddly,* SWE-2 is there for me.

If this performs similarly in the real world, we're approaching a level of capability where for most devs, it only makes sense to pay for Anthropic or OpenAI subscriptions if they are heavily subsidized and actually cheaper than these alternative options.

* Oddly, because I perceived Devin as being kind of a joke before trying SWE-2.

25m agoHN ↗

In their first video, to make a 3D render of the photos, the thinking traces gives away the game:

Interesting! It turns out there's already an existing project here [...] The project is fully built [...]

I'm always astounded how little effort is put into the details of these announcements. Back when I paid more attention, I remember OpenAI's and Google's demos having basic facts wrong every time, which is even worse than what Step 5 is doing here.