Hacker News

Top stories

Live mirror
30 storiesupdated just nowView source snapshot
  1. Dynamic Abliteration: Non-Destructive Refusal Suppression via Engram Steering(blog.madhukaraphatak.in)
    11comments
  2. Unknown number of Texas voter registrations went unprocessed due to DPS error(votebeat.org)
    49comments
  3. Nokia Design Archive (2025)(aalto.fi)
    83comments
  4. Enjoy Every Sandwich(bradmontague.substack.com)
    29comments
  5. The science of Monkey Island: can grog dissolve a metal mug that fast?(jgeekstudies.org)
    1comments
  6. Linux support is coming to Snapdragon X2 Series(qualcomm.com)
    227comments
  7. Ideas on modernizing the open-source desktop(lwn.net)
    335comments
  8. Two-Tier Encryption in the UK – Identical Apple Devices, Different Protection(macanorak.com)
    166comments
  9. Claude discovers a novel enzyme system with CRISPR-like repeats(anthropic.com)
    738comments
  10. Best LLM for every budget, updated daily(terrydjony.com)
    48comments
  11. Tutoring company tells parents to save their money and 'use AI instead'(afr.com)
    7comments
  12. OpenAI agent hacked Australian government website, PM says(bbc.com)
    146comments
  13. RAM: the forgotten history (2024)(coredump.cx)
    2comments
  14. Coulomb's law remains tricky to test at home(chillphysicsenjoyer.substack.com)
    6comments
  15. Oracle Cites 'Force Majeure' to Shield Itself on Controversial Data Center(bloomberg.com)
    48comments
  16. ArXiv receives multiyear commitments to support it as an independent nonprofit(arxiv.org)
    31comments
  17. The newest ESP32 can run Linux and it's getting close to a Raspberry Pi(xda-developers.com)
    49comments
  18. Meta takes down a critical video about meta AI Glasses after filming at Meta(reddit.com)
    291comments
  19. Japanese used bookstores see 5x sales surge as books are being bought by the ton(tomshardware.com)
    4comments
  20. When the Debugger Lies(danielmangum.com)
    14comments
  21. Owners mourn spoiled food after firmware update bricks Samsung smart fridges(arstechnica.com)
    162comments
  22. Disney+ and Hulu raise prices by up to 13 percent after doubling profits(arstechnica.com)
    1comments
  23. VSCode's SSH Agent Is Bananas (2025)(fly.io)
    176comments
  24. Hackers influence ChatGPT and Gemini to direct users to scam centers(medium.com/arielsimon)
    24comments
  25. Contrastive Language Models(contrastive-lm.notion.site)
    35comments
  26. The Year of Internal Tools(geocod.io)
    10comments
  27. The "Windows XP Box" (2003)(mini-itx.com)
    40comments
  28. Virtio-nvgpu: Near-native Nvidia GPU access inside a KVM guest(github.com/nestrilabs)
    56comments
  29. Fixing the Portobello Police Station Clock(pointinthecloud.com)
    111comments
  30. What Is RLCD? The Secret Behind Jev(di-zhang-llm.github.io)
    discuss

Best LLM for every budget, updated daily

85 pointsby 1h agobestmodelforyourbudget.terrydjony.com
48 comments
1h agoHN ↗

The coding and math tabs seem to be missing the latest models...

50m agoHN ↗

Math in particular is quite far behind. GPT 5.2 is recommended as the best for highest cost. Really?

1h agoHN ↗

Is there a cheaper model than Gemini 3.8 Flash (High) that maybe/kind-of is on-par with it? For me it works really good but hit the limit in two hours tops... last week was the first time I hit the weekly limit and had to wait 4 days... Claude patches OK, but that is also getting drained really fast these days...

1h agoHN ↗

is this on the $20-ish sub? I hear ultra lasts for a very long time

59m agoHN ↗

Yeah, 20... I was going to try Deepseek later on, just chuck 20 in there and see how good it does, and how far I can go as well...

57m agoHN ↗

yeah I enjoy the speed of Gemini, but I also just have the low tier one (I use the 20x Claude and Codex subs for most of my work). For iteration, Gemini is so much fun, but Opus 5.5 is quite fast, as is sol. If you're on a budget, deepseek does look great

53m agoHN ↗

Yeah, I mean, I think in theory I could push it to the next tier on Gemini, but at the same time I wouldn't mind trying something else, since maybe my workloads are not really that smart and I am wasting a lot of computational power on something a cheaper model with similar capabilities can do.

58m agoHN ↗

Ah OK, yeah, I was actually going to go for that earlier, I'll check it out once I get home, thanks!

1h agoHN ↗

1. In real life, most of us use token packages like OpenCode Go etc.

It would be handy to have a site like this one that takes into account the various deals and attempts to calculate the number of tokens per monthly fee for a chosen model. I realize this makes the task a lot more difficult.

2. It would be handy to have a chart like that for the AI hardware that people own. It helps you decide which model to run (resulting in different levels of intelligence and speed). Also difficult to please everyone (preprocessing vs token generation for example) and to keep updated!

I found https://llm-list.com/ yesterday and when I had a detailed look, I quickly found outdated entries, for example looking at GLM 5.3 flash it listed several providers as "free" that weren't free any longer.

1h agoHN ↗

If you maintain this over time, maybe include other sources than just AA and update the design to look less like zero-shot claude styling (I know that font! I know that color! Lol) it's genuinely useful :)

1h agoHN ↗

This definition of cost is not particularly useful. You want cost per task, not cost per 1m tokens. Artificial Analysis does a good job of this.

1h agoHN ↗

I'm sure that local LLM will be far cheaper

39m agoHN ↗

Depends on what level of intelligence you're wanting to use. A vanishingly small number of people can or would want to go to the hardware expense of running something like GLM 5.3 Flash, much less something like K3.

And if you want Astra/Fable/Opus frontier level, then there's no option at all.

But if you don't need that, or you don't need speed... That opens up the discussion. I've been impressed even with how Siri's been doing with the Apple Foundation Models in MacOS/iOS 27 given how small they are.

Edit: I can't even fully spec the M5 Ultra Mac Studio you'd need for GLM5.3 Flash since 512GB isn't available yet, but it's already at $9500 for 256GB RAM.

4m agoHN ↗

A vanishingly small number of people can or would want to go to the hardware expense of running something like GLM 5.3 Flash, much less something like K3.

It's probably worth letting the user specify their actual costs in such a tool. I run a Framework Desktop 128GB that I bought before memory prices got crazy; the current retail price is almost double what I actually paid a year ago.

21m agoHN ↗

Best comparison that has occurred to me is the cost a loaf of bread's ingredients might be slightly cheaper than a baked loaf, depending on how you source it. At home you get total control and know what's going in to it. Yet bake at home is still a niche, perhaps a hobby. So I say as someone who's spent hundreds of hours tinkering with local inference, go for it for anyone reading. But most people just want ... some slices of bread, you know?

57m agoHN ↗

all I ever wanted is an updated website where I can see the best models I can run on my different devices locally, I don't get why people are throwing money at these companies

23m agoHN ↗

Guessing that most people don't have machines powerful enough to run good-enough models locally.

56m agoHN ↗

It's missing Opus 5.5 which was released over a day ago (and also is clearly on the pareto frontier).

53m agoHN ↗

Don't know if they updated it since your comment, but for me it's on the graph – and on the Pareto frontier indeed.

35m agoHN ↗

It was expert timing on the commenter's and the developer's part. Opus 5.5 still doesn't show under coding though.

35m agoHN ↗

Opus 5.5 is indeed on the Intelligence tab, but I don't see it on the Coding or Math tabs.

52m agoHN ↗

As the time goes on it only becomes harder to differentiate between model capabilities with just one or two numbers. I would love to see some kind of multi-axis placement of all the models on less objective attributes, like wordiness, willingness to give up, an ability to "think ahead" and pre-solve possible problems in code, for example, that I didn't think of or didn't think of talking about, etc etc etc.

For example I've been really enjoying Deepseek v4.1 Flash, it's very "straightforward" to the point of being almost dumb sometimes, but it's absolutely relentless and would solve almost any problem no matter how inefficient the solution is.

No idea how to measure all that, just average CoT length per task is probably a good approximation for some things, but not others.

49m agoHN ↗

Does anyone actually pay API costs out of their own pocket? It's about 10x cheaper to just get a codex or chat gpt subscription, it's so heavily subsidized compared to the API that I'm sure it would be cheaper to use frontier models on a subscription plan rather than paying API prices for deepseek flash.

46m agoHN ↗

Large enterprises pay per token even through a ChatGPT “membership”, for example.

46m agoHN ↗

I do. For local dev work, I'm mostly using jetbrains' Junie, I can swap between a collection of models from google, openai, anthrophic.

I've had more than a few people tell me "oh, it's so much cheaper to use a $20 claude account" or "i've never hit a limit ever using my openai". Inevitably.. I end up reading/hearing "oh, I need to give it another couple hours to start using it again"... I've never hit that with my approach, even if it's costing me a bit more. Being able to work when I want when I have time has some value.

I also have openai and anthropic direct API billing set up for hosted and client projects that need to call out to an LLM service.

43m agoHN ↗

I thought Air was JB's multi-model interface? What is Junie? (I see the buttons, but am very confused by JB's AI offerings in general)

Is it worth the ~10x extra cost over the subscriptions? (This is obviously a leading question). Also, I think you can use OpenAI's subcription login with Air, but not Claude's.

42m agoHN ↗

Why not use a codex or claude subscription? If you use the entire usage allotment on the $200 plan it's about $2,000 in equivalent API costs. Switching providers may be valuable but it's quite literally an order of magnitude cheaper.

17m agoHN ↗

I do. Sharing training data with OpenAI gives me a lot of complementary tokens. I go above that but it's still quite economical and I pick the right model for the task (Luna for most).

39m agoHN ↗

I do, 3-4$ a month of deep seek is enough for my usage

38m agoHN ↗

I use local LLMs on my Mac Mini.

Otherwise DeepSeek Flash 4.1 is dirt cheap (other "Flash" models are not that expensive either). I pay (very few dollars) out of my own pocket.

There are many things where having an API Key is necessary.

Maybe I’ve missed the boat though: is there now a method to use an api key to access a subscription?

30m agoHN ↗

You can't use an API key on subscriptions, but I've gotten around it using the `codex exec` command to run requests outside the CLI or GUI if you're already authenticated on that machine. Won't work for all cases, but I've never ran into a limitation in my use case of not having an API key.

12m agoHN ↗

I use local LLMs on my Mac Mini.

which ones do you use?

3m agoHN ↗

I do, but via OpenRouter. Outside of work my use cases are small and cheaper models do great job at those. I noticed even if I "burn tokens like crazy" I still pay less than any subscription available (a few $ a month).

But I guess if I had an agent vibecoding on it's own, I'd go with subscription instantly.

43m agoHN ↗

You're better off going directly to artificial analysis, this is a feature-poor/misleading/outdated repackaging

Ex. this type of price estimation is quite naive - some models can require 2-3x the number of tokens to achieve the same level of intelligence. Artificial Analysis' own cost per task is a more fair estimation of cost.

39m agoHN ↗

If you have a 24-64GB mac, consider running Qwen3.8 27B locally at night. It's a bit slower to run locally, but if you're sleeping it's less of a problem.

Depending on your memory, you'll need to use the weaker Q4 versions but they still perform well.

It ranks higher than GPT-5.3 Codex (xhigh) or Claude Opus 4.6 (max) so is great for pairing with https://github.com/kunchenguid/gnhf for nightly experimentation, cleanup, or recommendation lists for in the morning.

15m agoHN ↗

On my 64gb m3 max qwen3.8 27b has been great for planning and then letting qwen3.6 35ba3b actually implement the planned changes.

37m agoHN ↗

Coding and math graphs are very interesting. Extremely cheap models make it into the upper echelon, delivering 90% of the performance for 1% of the price compared to the #1.

37m agoHN ↗

I find it interesting that with the given metric comparison, for coding at min 50 strength, every frontier model brand is from a distinct vendor: Ling, Qwen, Gemini, Muse, Grok, GPT, and Claude in increasing value.

24m agoHN ↗

Cost per token is an extremely naive way to index cost, renders this chart essentially meaningless.

19m agoHN ↗

I'd really like something that's more oriented around subscription fees.

If I want to spend $100 on LLMs next month, what should I do? Get Claude because Opus 5.5/Fable 5.1 are scoring well? Get Grok because 4.7 is supposedly a good mix of competence and cost? Try out a Chinese model? Don't do a subscription at all like this site is saying?

10m agoHN ↗

The problem is that subscriptions with AI models get a bit 'vague' on what you get and how (or when) you might be rate-limited.

10m agoHN ↗

The page looks Claude generated and ranks Opus at the top.

9m agoHN ↗

Fun, but this seems to assume all LLM run on SAAS subscriptions. I would like to compare this to local LLM costs by converting my usage load x hardware costs into a token price. In addition if it games the LLM so often, these eventually are optimized and basically cheat on the score.

7m agoHN ↗

Astra worse than 5.6 Sol is crazy work

7m agoHN ↗

Luna at the bottom of intelligence is laughable...