Hacker News

New stories

Live mirror
30 storiesupdated just nowView source snapshot
  1. Anthropic CEO Amodei set to meet with Trump (cnbc.com)
    —discuss
  2. HiFinder, a map for outdoor spots that never sends photos to a third party (hifinder.net)
    —discuss
  3. Debt-hungry AI companies face increased risk as bond yields spike (cnbc.com)
    —discuss
  4. Allegations of US interference in Quebec election (globalnews.ca)
    —discuss
  5. Show HN: Kernal Fit – how I lost weight working IT without a gym (adrenaline.design)
    —discuss
  6. Show HN: Cartopolis, interactive globe-sized 3D world (garage44.eu)
    —discuss
  7. Luarocks.org remote code execution exploit (neorg.org)
    —discuss
  8. Booted up in 1993, this server still runs – but not for much longer (2017) (computerworld.com)
    —discuss
  9. Walmart says it's not using personal information to set prices (apnews.com)
    —discuss
  10. When did Google get so f-ing weird? (sancho.bearblog.dev)
    —discuss
  11. The largest privately owned laser just turned on (techcrunch.com)
    —discuss
  12. A year of hourly weather drawn as the iris of an eye (experiai.com)
    —discuss
  13. Montessori Education (wikipedia.org)
    —discuss
  14. Stake data breach exposes customer details after hack at partner DriveWealth (news.com.au)
    —discuss
  15. Have an LLC (zachholman.com)
    —discuss
  16. A crypto sentiment API that AI agents pay for autonomously (x402 protocol) (crypto-sentiment-x402.onrender.com)
    —discuss
  17. Self-Hosting on the Dark Web (alvarezrosa.com)
    —discuss
  18. Why Copy and Paste Still Feels Wrong on a Phone (medium.com/ai-widgets)
    —discuss
  19. Design Decision: Technical Debt in BillaBear (iain.rocks)
    —discuss
  20. My first time in Stockholm (thanks, YC) (domelian.substack.com)
    1comments
  21. Oil Slips as Iran Diplomacy Sparks Fresh Hopes for a Market Reset (medium.com/d02514047)
    2comments
  22. Preface – Home Manager Manual (nix-community.github.io)
    —discuss
  23. Introduction – Act – User Guide – Manual – Docs – Documentation (nektosact.com)
    —discuss
  24. Show HN: An AI wedding concierge I built and ran at my own 300-person wedding (ai-do.io)
    —discuss
  25. US push to flood Balkans with American gas sparks concerns in Brussels (politico.eu)
    —discuss
  26. An OpenAl Engineer and His Friends Debate the Future [video] (youtube.com)
    —discuss
  27. Corporate America embraces cheaper 'open' AI models (ft.com)
    —discuss
  28. TuxBot: Semantic-Aware Online OS Tuning with Large Language Models (arxiv.org)
    —discuss
  29. DSPy – Program, don't prompt, your LLMs (dspy.ai)
    1comments
  30. Why YC Companies Win (khaliqgant.com)
    —discuss

Ember-1

173 pointsby 2h agofireworks.ai
95 comments
2h agoHN ↗

The problem: thinking models think too much

Analysis paralysis stifles not just human intelligence, but other intelligences too.

1h agoHN ↗

Yes and thar makes you wonder if the Paradox of Choice would apply as well ;)

The more options you have, the harder it becomes to be satisfied with the one you picked.

54m agoHN ↗

The thinking traces on some Chinese models just output the full response in the thinking trace, then output it again to the user, which is redundant.

2h agoHN ↗

I don't think the article mentions Pareto frontier enough.

Also, did I miss a memo? Suddenly every article on AI seems to be talking about the Pareto frontier - or have I just not been paying attention?

1h agoHN ↗

I guess they figure "best bang for your buck" comes off a little too colloquial.

9m agoHN ↗

I would really love if we brought back some colloquialisms in this field. Not that long ago most folks in tech would have had pretty blank looks on their faces when someone started talking about the "Pareto frontier"

1h agoHN ↗

They want it to be the best at something. And it's obviously not the absolute smartest. So here we are.

1h agoHN ↗

Pareto frontier on some benchmark that I am hearing of for the first time.

Kimi K3 with less reasoning tokens isn't exactly exciting either, and particularly so if the license is less open than original Kimi K3.

4m agoHN ↗

when everyones fighting to be 'somewhere in the pile' they need some way to advertise they have made progress while not being the best.

2h agoHN ↗

Need this done for DeepSeek, ideally one of the Flash models.

1h agoHN ↗

If you have the compute, I have the expertise.

1h agoHN ↗

And GLM. Both Deepseek 4.1 Flash and GLM 5.3 Flash are quote verbose when thinking.

2h agoHN ↗

Well done, and great iteration.

The pareto frontier needs clearer distinction. Benchmarks miss half the story. What, if any, capability is lost by the token reduction (for example, was it like super awesome at Golang before and now kind of sucks? that kind of distinction).

1h agoHN ↗

Unfortunately, it’s hard to make a chart of that.

1h agoHN ↗

It looks like it would be similar to GLM 5.3 Flash, had they tested it...

1h agoHN ↗

Ignoring for the moment issues of what "counts" as open, won't open models rapidly advance due to stuff like this in ways that it's less possible for the proprietary ones to do? This is exactly how Linux & Wikipedia, for example, overtook their "frontiers", right?

1h agoHN ↗

Yes, absolutely, but only if people keep contributing in the open.

1h agoHN ↗

not necessarily, just knowing something is possible will motivate others to achieve it somehow. Which is why there are so many LLMs and OAI doesn't have a monopoly

1h agoHN ↗

No, because close labs/models borrow but don't contribute back.

1h agoHN ↗

Ignoring for the moment issues of what "counts" as open, won't open models rapidly advance due to stuff like this in ways that it's less possible for the proprietary ones to do? This is exactly how Linux & Wikipedia, for example, overtook their "frontiers", right?

I suspect the advantage that catapulted Linux ahead of the establishment was less technical potential and talent and more organizational advantage. That's not to diminish the technical talent of the Linux crew, but them being unencumbered gave them more degrees of freedom. The rest is history.

So as long as the AI companies don't succumb to "big company" dynamics, they can outlead. To wit: Open AI and Anthropic are kicking Google's ass.

1h agoHN ↗

Diff people have diff motives to experiment, then new work is done on top of stuff that "hits" in a way no one anticipated. Then work gets piled on top in a way that might make it hard to port

1h agoHN ↗

then new work is done on top of stuff that "hits" in a way no one anticipated.

Indeed. And when you have freedom to play, you are able to find new stepping stones that you didn't anticipate. And you can combine stepping stones in new ways to make new discoveries.

Greatness cannot be planned.

52m agoHN ↗

I think people make mistake here, google’s approach is not to spend $2.3 on every $1.0 earned, they’re riding on serving to masses “luna”, they absolutely have way more powerful models internally but they don’t clutter their infrastructure with fragile and costly intelligence-of-size inference frontier. I think “underdog” perception is illusory/temporary, not stupidity - calculated, conscious, longer term bet.

1h agoHN ↗

The difference between contributing to OS and AI, is that the first is a hobby alternative to woodworking or hiking, while the other can easily bootstrap you a company you can get millions in investment, at least for time being.

1h agoHN ↗

Won't the "frontier" labs figure out whatever techniques were used and apply them to their closed models?

1h agoHN ↗

If they can keep up.

The lock-in is less pronounced as it is with AWS or MS.

1h agoHN ↗

Like how the last 2 decades of tech companies are thinly veiled open source pilfering into business units.

1h agoHN ↗

On the smaller end, Quen 3.8, while being extraordinarily capable for a small local model, also suffers from extreme thinking. I wonder if the techniques described here generalize to other models too.

1h agoHN ↗

I suspect it might generalize to other large models, but I don't think Qwen3.8 27B is one of them. Kimi K3 is a 2.8 trillion parameter model, and I suspect that is playing a big role in being able to reduce the length of CoT without taking a hit in quality.

That's just vibes, though.

1h agoHN ↗

Does anybody know if this would be a good model for creative writing?

1h agoHN ↗

So they trained a model on open weights, and then aren't releasing the weights... am I reading this right?

1h agoHN ↗

Technically kimi k-3 weights license is not open weight (it has a lot of restrictions). I would classify it as ‘weight open’ similar to the bsl and fsl ’source open’ licenses.

1h agoHN ↗

It happens. Most open licenses aren't GPL style copyleft.

1h agoHN ↗

It happens with open source software all the time, why would we expect any different with open source weights.

1h agoHN ↗

Because we do. The GPL isn't a suggestion. If you can take open source code and make private software out of it then what are we all doing? No, license requirements and agreement are law for a reason.

4m agoHN ↗

GPL is a specific license, it’s not FLOSS as a whole

1h agoHN ↗

Because the licenses that apply to software make no sense in the context of LLMs. With the latter, there is no source code to license.

The words of a license are what the license is.

1h agoHN ↗

Aren't Cursor Composer models like this too? At some point all the extra RL you do can be considered as proprietary information added.

Not suggesting this is right or wrong, but is sort of the nature of the technology.

1h agoHN ↗

There is little to no point reading the article as well. It's stripped of all alpha.

task and environment feedback

on-policy planning and learning

feedback connects decisions to their consequences

These are deliberately the least informative phrases you could possibly use to describe what you have done, while still being in the realm of words that go over a generic investor who has no idea whats going on and may be dazzled by sciencey sounding language.

Cursor compose 2.5 article where they used and described on policy self distilation was actual alpha.

5m agoHN ↗

Which is fine, that’s legal according to the license

1h agoHN ↗

Off topic:With sol pricing drop tbh kimi k3’s value prop has not been that great. For our internal use case/testing/benchmarks sol come out with way better quality and much cheaper costs. Kimi really needs to drop their pricing (I heard it’s set by them across all the neoclouds) Sol is at 2/10 vs kimi’s 3/15

1h agoHN ↗

Agreed. Even on the open weight side, GLM 5.3 has roughly equivalent performance to Kimi K3 for less than half the cost.

1h agoHN ↗

Agreed, I think the only place where it’s still interesting is ui design. Visually kimi and muse feel much nicer than frontier models to me, but maybe it’s an artifact of everything terrible being Claude Design

49m agoHN ↗

Competition is good. Without K3/GLM/DS4 etc. there would be no pressure on OpenAI to drop Sol's price.

18m agoHN ↗

I was surprised by that. I run my benchmark [1] every couple of days and was sure this model will be ath the pareto frontier, if not THE pareto frontier. But no:

Ember isn't picked yet. In planning, Opus 5.5 wins under the planning weights. In code, GPT-6 Sol dominates it: also 10/10, but with a higher quality score and a lower estimated cost. Ember has no intelligence index, so its starting score is only 0.73, which holds its 10/10 down to 0.954 against Sol's 0.975.

[1] https://philippdubach.com/posts/jev-model-router-for-pi/

15m agoHN ↗

With sol pricing drop

6 or 5.6? Because 6 is hot garbage

1h agoHN ↗

This is really interesting. I think the Fireworks Serverless Training infrastructure they used to develop it is also unique and needed. Except if someone works at one of a handful of the largest labs, it is very difficult to set up or try any sort of training pipeline. The managed training infrastructure makes it available to more people.

1h agoHN ↗

I can’t help but think it’s more expensive tinker.

1h agoHN ↗

The problem: thinking models think too much

This is partly the appeal of Jev et al; having a quick model for simple tasks, that doesn’t require that much thinking

It’s amazing all the workflows that models like that can unlock. And yes, classifiers and other ML models have been around for a while for these types of tasks, but Jev has made it easy and cheap to play and experiment. This in turn, is incentivizing people to try them for a bunch of stuff, unlocking creativity and producing a lot of new cool (and eventually potentially very useful) applications

1h agoHN ↗

What are the useful applications of Jev so far? Not to sound dismissive, I just haven’t seen what people are using it for yet.

1h agoHN ↗

Lots of use cases! I've personally used it for the following:

1. Evals (once you have your rubric defined and tuned using a reasoning model, jev can be great for running periodic evals especially those that run daily.

2. e-commerce catalog classification 3. quick search using anything as context and query mapping to a pre-defined set.

8m agoHN ↗

At least for 1, evils, you’d want to use a good old reasoning model to get the best eval results.

1h agoHN ↗

Why not using a cheap LLM with thinking completely disabled ? I don't think it will be much more expensive than jev.

56m agoHN ↗

I’ve tested this with some local LLMs and their accuracy is in general better than Jev/Laya, but they are super slow in comparison as well

For example, a typical/stock LLM can’t really play Doom in real time, but a Jev-like model can. Just because of latency

Of course, if you want the best Doom player, there are way better and faster adhoc models

33m agoHN ↗

LLM inference has two very different regimes of work: prefill & decode. You can think of the former roughly as processing a pre-specified prompt, and the latter as sequential processing (auto-regressive token generation) eg. "chain of thought". The latter is very important for LLMs and cannot be ignored; it deeply influences infra design, even necessitates copious amounts of high-bandwidth memory. Jev-like models can ignore the latter and therefore optimize much better for the former, consequently operating at both better cost and latency.

1h agoHN ↗

The result? Ember-1 set a new Pareto frontier for Bedside Bench across both open and closed models including GPT-5.6 Sol, GPT-6 Astra, and Claude Opus 5 on cost/task.

"Pareto": 8 hits

"Opus 5.5": zero hits

1h agoHN ↗

Obviously this research was done before 6.0 Sol and Opus 5.5 came out. Your point stands that the frontier moves quickly and small gains can be eclipsed quickly.

1h agoHN ↗

This is the golden age of model training. Some days ago, I decided I wanted a local CPU only model that can perform exceptionally well for English to Bash translation (to avoid the googling for command syntax). I got a bunch of subagents to generate large amount of training data (140k+ samples), got the Qwen 3 0.6B base model, pointed Astra at it, and off to the races. It trained for 2 days (on and off) and I got a surprisingly good model for my task! The total active time I spent was a few hours. And it is still improving, what a time to be alive!

1h agoHN ↗

I didn't have a local GPU, so I asked it to go out and find hardware. It found a google TPU v6e which seemed reasonably priced. I gave it my google api key. I told it to use TPU only when training and bring it down afterwards. That's about it.

1h agoHN ↗

What kind of observability did you have over this process? I’m interested in how my peers are operating these efforts.

52m agoHN ↗

On the cloud side, nothing valuable existed, so the training couldn't ruin anything it didn't create. On the laptop side, I usually ask the agents to create named scripts for everything it needs to access, then those local script directory is green-lit with approve all. For cost, I kept giving it new budget in the 20-30 dollar increments.

I had to intervene a few times. For instance, as smart as the models are said to be (Astra), it would copy the full training run, train on the server, pull every checkpoint to the local machine, then run tests, update. So, the bandwidth bill was as high as training bill for the first 6 hours. It could have simply tested each checkpoint on the server, saved time and money, didn't occur to it until I said.

47m agoHN ↗

I told it to use TPU only when training and bring it down afterwards.

I wouldn't put my house on it. Brave.

41m agoHN ↗

I gave it my google api key

This is the part where the narrator looks at the camera and says "Don't try this at home, kids!"

35m agoHN ↗

Why? Isnt the API key scoped to a project and specifically made for this?

Are you confusing this with an OAuth token or something?

18m agoHN ↗

Until astra goes bonkers and use the tpu for days

34m agoHN ↗

There’s a safer way to do this with nearly no added friction. Give it a read only API key. Then just ask it to write the API calls into a bash script and then read it and run it yourself. The agent can still inspect the live resources and diagnose and give you more commands to run. I do agree I wouldn’t give it create / write access.

33m agoHN ↗

You’re absolutely right, I shouldn’t have rented a 200 GPU cluster for $35,000/hour. That’s on me.

[Search: Can I refund Google cloud?]

It looks like we’re not able to ask for a refund since we did actually use all of that compute intentionally.

Would you like me to write you a pleading email to send to the support team?

28m agoHN ↗

I've done this sort of thing before but with Vast. Pre-deposited some money online, then let the LLM request and manage a training run on an allocation. Worked pretty well without risking bankruptcy.

1h agoHN ↗

I am thinking about opensourcing everything, although this is not my main domain or my main startup, so the overhead of huggingface etc seems a bit unnecessary

Edit: will do as soon as possible

53m agoHN ↗

just ask the agent to write it up if you don't have time to do a write-up yourself

35m agoHN ↗

+1, would like to see. Even if it's not fully "ready for consumption", it's probably enough to reproduce the results.

26m agoHN ↗

Please do! Small, specialized models need more love and the time you spent would be a gift!

25m agoHN ↗

Would also love to read a write-up about this!

48m agoHN ↗

Curious about how you generated the training data? Was it just asking an existing model to generate a bunch of examples?

I ask cause would this be a kind of model distillation?

I have a small model I'm looking to train on some data, and I have some real live data but I'd love to be able to extend it.

37m agoHN ↗

All synthetic data. For this usecase, it was easier because all current generation LLMs, even the small models, are really good at bash commands (and SQL queries too)), so you can reasonably start batches of cheap subagents whose output is reviewed by a more capable model and merge into main training set. After 100k, I had to standing instructions to run the generation loops selectively, meaning only update samples in a given area where we see poor capability.

16m agoHN ↗

It is a form of distillation, as long as you're working a very narrow "trivial" topics it works perfectly.

36m agoHN ↗

That's a really impressive result. There are all kinds of small tasks like this I use an LLM for, but theoretically if you broke all the sub-use cases into local-only models, and had something lightweight that routed to the right model, you could have faster and cheaper workflows. E.g. something trained on the linux man pages for common commands, since it's usually quicker to ask an LLM for a specific command with flags than to consult the man pages.

19m agoHN ↗

I don't understand. If you have a model that can do bash examples already (your subagents), then why would you need to train a model?

Or are the subagents generating your training data using a closed/paid model?

13m agoHN ↗

The models he is using to generate training data are presumably commercial models. He is distilling their bash knowledge into a much smaller model he can run locally fast and cheap.

12m agoHN ↗

A very small, highly specialized model can use negligible resources (CPU, energy) to accomplish the same task.

For everyday work that happens frequently it's better to have a tiny specialized model instead of making billable API calls or turning your laptop into an 80W space heater for 20 seconds to run a general purpose model.

The large models can be used to generate synthetic training data. Tell them to make up 100,000 tasks paired with the resulting output as a 1-time cost. Then use that to train a small model.

Think of it as distillation, but focused on a specific task.

4m agoHN ↗

If it’s one of thing that you want just for English to bash shell commands, I will create AST, it is deterministic, exceptionally fast, no tokens so no need to fine tune existing model, please let me know your thoughts.

39m agoHN ↗

Aside. I find the "cost per task" charts both useful and uncanny. Is It better a model that takes me to 90% in 1 dollar or one that takes me to 95% in 2 dollars? Or a different model that too scores 90% in 1 dollar? How much will it cost me the last 10% or 5%? At the end of the day, cost to 100% is what matters and the half (90%) backed solution may require more to reach 100% (or not, who knows?)

11m agoHN ↗

Is It better a model that takes me to 90% in 1 dollar or one that takes me to 95% in 2 dollars?

It's pretty important to understand if your own work domain is one where the last 5% matters. In a lot of day-to-day software engineering tasks, it doesn't, and one can get crazy mileage out of the cheaper models. OTOH, if you are performing novel research, that last 5% may be worth whatever it costs...

30m agoHN ↗

The problem: thinking models think too much

I see that with Opus 5, it started thinking like crazy in the last few days , I don't think my workflow is that complicated, still it gets into thinking mode and stays there