Hacker News

Best stories

Live mirror
30 storiesupdated just nowView source snapshot
  1. When did Google get so weird? (sancho.bearblog.dev)
    1052comments
  2. Owed a billion dollars in Nvidia stock (colo.to)
    447comments
  3. Sonnet 5.5 (anthropic.com)
    492comments
  4. Unsealed Briefs in Authors’ Case v. Microsoft/OpenAI (authorsguild.org)
    612comments
  5. Ember-1 (fireworks.ai)
    247comments
  6. Pirating the Pirates (mubi.com)
    256comments
  7. Coding is not solved (alexewerlof.com)
    476comments
  8. Windows 11½ (definitelynotwindows.com)
    144comments
  9. Meta Blocks President Lula's Facebook Page, Campaign Ads 2 Weeks from Election (reddit.com)
    303comments
  10. Jeff – Jev-compatible 0.8B decision models, trained at home, ~30 ms (github.com/firelex)
    168comments
  11. Updated Google Maps shows destruction of the city of Rafah (twitter.com/aliabunimah)
    245comments
  12. It's Time to Investigate the AI Labs (calnewport.com)
    163comments
  13. AI companies in race to demonstrate their model most threatening to humanity (thecivilian.co.nz)
    385comments
  14. There are no "rogue" AI agents (eoinhiggins.substack.com)
    268comments
  15. On caring for user data: NeoVim caused Vim undo files to be deleted (aresluna.org)
    343comments
  16. Tells of a Slop UI (hereticpleb.vercel.app)
    241comments
  17. Kids turned low-traffic NPR Spotify comments into a secret group chat (thisamericanlife.org)
    207comments
  18. The problem is not AI code, but not knowing about system architecture or intent (ssp.sh)
    227comments
  19. MongoDB CEO resigns to join Meta (reuters.com)
    280comments
  20. Self-Hosting on the Dark Web (alvarezrosa.com)
    113comments
  21. Don't couple your Go code to GitHub (iain.rocks)
    175comments
  22. Parley: Federated, decentralised chat that speaks plain IRC (mills.io)
    175comments
  23. Show HN: Lofi Cities – Pixel-art city nights with browser-generated lofi (loficities.com)
    135comments
  24. SpaceX's Starship launching to orbit for first time ever today (space.com)
    341comments
  25. In an $80 motel room, a discovery to shed light on the origins of life (nytimes.com)
    106comments
  26. The Normalization of Inexplicable Failures (ihatethefuture.com)
    123comments
  27. World Labs Is Joining AMD (worldlabs.ai)
    108comments
  28. Hijacking the PS5's RTMP stream (yashgarg.dev)
    75comments
  29. So long Google, and thanks for all the nudes (lecaro.me)
    90comments
  30. MicroLLM Lab – Try 7 tiny LLM's in the browser (stateofutopia.com)
    77comments

MicroLLM Lab – Try 7 tiny LLM's in the browser

212 pointsby 12h agostateofutopia.com
77 comments
12h agoHN ↗

You can try 7 different tiny LLM's in your browser.

11h agoHN ↗

What I missed from the title was this: Can I try 7 different tiny LLMs in my browser?

11h agoHN ↗

i have a bromine tub, give me simple instructions for what to do since i just filled it up w fresh water

1. Add 1 tablespoon of water to the water bath. 2. Place the tub into the water bath and let it sit for about 5 minutes. 3. After 5 minutes, remove the tub and let it cool down. 4. Now, fill the tub with water and let it sit for about

uhmm completely unusable ?

10h agoHN ↗

No, that is not what SLMs are useful for. They have very little knowledge, and typically lack reasoning skills.

Useful applications include sentiment analysis, text classification, entity extraction, etc.

They certainly can be useful, but you shouldn't compare them with LLMs such as Opus or Fable.

10h agoHN ↗

Completely unusable for that, yeah, that's a 100M model, not an assistant

11h agoHN ↗

The website literally says they are SLM's .. derp af

11h agoHN ↗

I did the default arithmetic with PetitGPT research-v1

What is 2+2?

Answer

To find 2 + 2, we need to add 2 to both sides of the equation.

2 + 2 = 4

So, 2 + 2 = 4 + 2.

Brilliant

9h agoHN ↗

That's not incorrect.

LLMs produce semantically correct sentences, not factually correct statements. Have we forgotten this so soon?

9h agoHN ↗

It is, in every sense, incorrect. Which statement in this short snippet is "semantically correct"? (Better LLMs get this right, of course.)

8h agoHN ↗

It's not very good semantically either.

Give me a recipe for soup.

Here is a recipe for soup:

Saffa-Cake-Sweet-Sweet-Sweet-Sweet-Sweet-Sweet-Sweet-Sweet-Sweet-Sweet-Sweet-Sweet-Sweet-Sweet-Sweet-Sweet-

9h agoHN ↗

GPT-2 124M is terrible...

  > what is 2+2?

Answer:

  > 3+3? 4+4? 5+6? 7+8?
  > Reply ~18000 0 ~10 min 2 By : 1-1: I'm a beginner. 3x2 is my best option, but if you're not sure about the other options then just go for it and try again
9h agoHN ↗

It's not instruct tuned looks like? It's closer to a base model rather than a chatbot.

8h agoHN ↗

That one is a 2019 model :) Years before the ChatGPT public preview.

8h agoHN ↗

GPT-2 is an interesting one because it is a February 2019 model. (You can see some information about it below the card if you click on the card.)

That was 2-3 years before the big "ChatGPT moment" (the highly coherent ChatGPT research preview was released in November 2022, I think it was ChatGPT 3.5). Back in 2019 the models really were not producing very coherent output. Now you can see it for yourself right in your browser :) Everything has come a really long way since then!

8h agoHN ↗

I've been playing with it for a bit, and I'm a tiny little bit surprised that they decided to continue pursuing that direction of research at all. What I'm getting from it looks like it could just be arbitrary sentence and paragraph fragments from the Internet, pasted together Markov-chain like.

I'm not sure I would have ever believed that something useful would come out of it, yet here we are.

5h agoHN ↗

Sounds about right about something that's 124M. I trained one myself a few weeks ago and it's about the same level of being terrible.

8h agoHN ↗

Awesome! I tried it and after loading (which took a while as it is a large model) got 12 tokens/second and very coherent output. Great demonstration.

7h agoHN ↗

I would recommend adding the MiniCPM5-1B model. Surprisingly coherent for a 1B model and performs well. On my pixel 9 I get 33 tok/s on CPU. On GPU I get about 26 tok/s but prefill jumps to nearly 500.

10h agoHN ↗

Cool project, but I'd really suggest looking at the UI.

The text is too small and it's way too dense with information in general. Considering how simple this product is to use, it's kinda crazy that I have to scroll through over a page length of (mostly useless, AI-generated) information before getting to the actual interface.

Also what is going on with the footer (why does it link back to the site itself, why is it telling me to "serve over HTTP").

9h agoHN ↗

This is the future of software, sloppy ui.

6h agoHN ↗

To be fair this is worse.

It’s verbose with a layer of looking legit play first glance sloppy UI.

I don’t mind vibe code ui but at least some ui efffort would be nice. Not just 1 shot.

6h agoHN ↗

I think your memory of pre-AI UIs is rosy. This is a little below average.

6h agoHN ↗

We didn't have this kind of sloppy UI. Putting in an extra text field is work, so we mostly had blocks of text and misaligned fields and text in the wrong place.

AI has no trouble churning out code, so AI-generated sloppy UIs have text fields everywhere.

9h agoHN ↗

Thank you for the feedback, I'll think about how to incorporate the changes you've suggested.

Edit: I've concluded that the text is tiny so people can skip it, it's like fine print. There are only 600 words on the entire page, that's not a lot even if someone does read all of them.

Since this submission has been on the front page of HN for hours now, clearly people like it as it is. I don't think I need to change anything about it.

The footer means that if you download the zip file and unzip it, you can't just open index.html and expect it to work, you need to start a local webserver for it.

7h agoHN ↗

It's clear that this was vibecoded. I'm antivibecoding, but here's a tip.

Don't focus on the actual code or the actual UI. You need to give your model a way to see its output. If it just pumps code, it looks at the code. If you let it look at the rendered UI, whether by manually taking screenshots, or making it so that the agent looks at the browser output automagically (with a webdriver or selenium or puppeteer or a skill or whatever), it will fix whatever issues it sees on itself, and it will never be content with delivering something that looks as bad as this.

6h agoHN ↗

Since this submission has been on the front page of HN for hours now, clearly people like it as it is. I don't think I need to change anything about it.

That's not really how it works. People like it in spite of the bad vibe coded UI. That doesn't mean that the UI is good. The criticism was clearly correct. I don't feel bad about publicly criticizing the UI to your face, either, because it's not like you were the one who designed/created it anyway--by the same reasoning, I don't see any reason to get defensive about it, since you have not invested very much of yourself into it either.

This has become a bit of a pet peeve of mine in general. Over the last six months or so I've been inundated with all kinds of lazy, bad vibe-coded UIs via my normal internet usage. It feels paradoxical in a way, because AI tooling should make it easier than ever to fix obvious problems, and yet UIs have gotten worse.

I have two theories for why this is: (1) usage of AI tooling conditions people into engaging with their creations in a certain hands-off, uninvolved, low-attention surface level way; or perhaps (2) those who care about their UIs were the people who were already creating UIs pre-AI, and so the marginal effect of AI tooling has been to increase UI creation amongst precisely the population who cares the least about what they are putting out.

I've also been developing a very strong personal revulsion towards the Claude default design language, including and especially the gold-on-black color scheme, although I admit that this is entirely by association and not because of its inherent qualities per se.

6h agoHN ↗

In this particular case it used to not have any text at the top. Just click a model, chat.

I decided that that wasn't very user friendly. It doesn't tell people what it is. Or how to use it. It doesn't explain any of the concepts or vocabulary.

So I asked it to add those things, above the main interface, including the steps people need to take to use it. I think it did great, and I think the small text sizes with a clear title let people who need it, read it, and people who don't can just scroll down. Even with the text it's 600 words.

I like the effect of what it's done. It's better than I would have done if I had written those intro words myself.

Clearly it resonates with people, this thing has been on the front page of HN for 4 hours and with 110 points is by far the most popular thing I've ever posted here.

5h agoHN ↗

It isn’t about how small the text is or how easy it is to skip. Humans are not transformers with multiheaded anttention. Human brains do not process one million tokens in parallel. Humans have caloric constraints and will be annoyed if you ask them to process irrelevant information.

Therefore, when showing something to humans, the first thing they see should be the first thing you want them to see. And since this is a demo, the first thing they see should be the demo, with sane defaults, above the fold.

Also, it is likely that no one (including the author) has read this “fine print” because it includes incorrect information. GPT-4 isn’t a frontier model anymore.

7h agoHN ↗

The text is too small and it's way too dense with information in general.

The Claude special.

5h agoHN ↗

For real. all these dashboards look quite similar

7h agoHN ↗

It's a cool idea. I'm on my phone now (going to sleep soon), a few of the demos didn't boot on this device. I'll check from desktop tomorrow.

1h agoHN ↗

I tried the Python and SQLite modules on desktop Chrome and got similar errors for both. For Python:

Error: CPython engine could not load (Failed to fetch dynamically imported module: https://sonistellar.com/lab/vendor/pyodide/pyodide.mjs). Check the network, then close this tile and re-boot to retry. — press reset or back to menu

For SQLite:

Error: SQLite engine could not load (GET https://sonistellar.com/lab/vendor/sqlite/sql-wasm.js -> HTTP 404). Check the network, then close this tile and re-boot to retry. — press reset or back to menu

FreeDOS loaded though, Snake and Tetris were fun!

10h agoHN ↗

PetitGPT told me that

"2+2 is 2."

Otherwise, a very neat demo. As others have said, the UI is VERY confusing, way too much stuff going on.

1h agoHN ↗

Using the same model, I asked it what 2 + 2 is and it said "2 + 2 is 2 + 2" which is not wrong.

Other questions and answers were mixed:

How many cards in a deck? > A deck of cards contains 52 cards.

How many cards in a deck if I remove all Queens from the deck? > If you remove all Queens from the deck, there are still 52 cards in the deck.

10h agoHN ↗

Really cool project. Giving web apps direct access to on-device models is something I’m excited about, and it’s cool to see the different approaches.

I’ve been working on a related proposal called the Web Models API, which explores a browser standard for an API that runs open-weight models on-device. Would love your thoughts: https://www.webmodels.dev

10h agoHN ↗

I read your proposal, I think it's great! Where will the navigator get the model if the user agrees to download it? For this demonstration I just serve the models on my own server, but for larger models it may be an issue as they may not have direct download links even if they are open weights.

10h agoHN ↗

Great question. Right now there isn’t a definitive answer, but it’s something that needs to be worked out. There would likely be a registry. The question is how to keep model IDs consistent: does each browser manage its own registry, or is there one shared across browsers?

9h agoHN ↗

Since you're asking for some kinds of permissions anyway, you could ask if the user is willing to also seed the model, p2p. (However, seeding files is not as popular as it used to be, many residential Internet connections don't have good upload.) If you have the capacity for it, your site webmodels.dev could act as a tracker and initial seed for any models. Then it could be the one central registry. It might get to be too much for you though, a lot of the open weights models are huge.

9h agoHN ↗

Interesting idea. I hadn't considered that. Thanks

10h agoHN ↗

unfortunately on firefox: Uncaught ReferenceError: GPUShaderStage is not defined

9h agoHN ↗

I tested it on Firefox on windows, version 156.0.1 and didn't get that error.

What version of Firefox are you using and what is your operating system and graphics card, please? Can you also try it without WebGPU? (Reload the page and uncheck "Prefer WebGPU" and try a prompt.)

7h agoHN ↗

Same issue. Latest version of Firefox 156.0 on Ubuntu GNOME. AMD Radeon 860M.

Unticking the WebGPU checkbox and reloading didn't do anything.

7h agoHN ↗

This is a difficult one for me to fix since I don't have a Radeon GPU. I'll see if there is anything I can do tomorrow.

4h agoHN ↗

I got the same error with an Intel Iris. I'm using Firefox on Fedora Gnome. "Prefer WebGPU" also didn't fix it.

8h agoHN ↗

It's a cool project, but I couldn't get it to load. What did you test this on? I tried the live link here:

https://willaaam.github.io/gemma-4-E2B-webgpu-vision/

And after loading it, with an NVidia 1060 GPU (6 GB RAM) on Windows it failed with "Failed to load: No supported WebGPU variant for com.xenova.gemma4.DenseGemv; rejected sgma".

In Safari on a 2026 Mac Mini M4 with 24 GB of RAM it failed with "Failed to load: JSON Parse error: Unexpected EOF".

The idea is pretty cool though!

7h agoHN ↗

Huh, weird! It doesn't work in Firefox, but I test on Safari (M5 Max) and Chrome (Linux, Arc B580). I don't have access to NVidia or AMD hardware myself unfortunately, though friends confirmed both as working.

Going to debug tomorrow!

9h agoHN ↗

What browser are you using? I tested it on Windows, Mac, and iPhone. I tested it in Chrome, Firefox, Edge, and Safari. Everything works on the three machines and phone I tested it on.

(It's a little bit slow at the moment - you have to wait a few seconds for the models to load - as it's currently on the HN front page. The server is on a 1 gigabit unmetered network connection so it can serve all the weights - around 600 megabytes - to one person every few seconds, there are several concurrent users now.)

9h agoHN ↗

how would you compare your capabilities to that of claude fable 5.1 by anthropic

Comparing your capabilities to that of claude fable 5.1 would be very similar. Both are stories about a clown who is a cunning and manipulative animal. They both have problems and challenges, and they face challenges to overcome them. However, they have different goals and approaches. The clown

I think that tells you all you need to know.

6h agoHN ↗

That's hilarious. At least the smaller model is more self aware!

8h agoHN ↗

Thanks. Unfortunately it's too late to change the title!

7h agoHN ↗

Maybe we should just call them SLMs instead of LLMs...

1h agoHN ↗

Interesting project, thanks for sharing.

5h agoHN ↗

Amusingly limited: I asked SmolLM2 360M Instruct

"What is the population of California?"

It replied

"California is the third-largest state in the United States, with a population of around 153 million people. It is also the most populous state in the United States, with a population of around 325 million people. The state is also the most populous in terms of GDP, with a GDP of $1"

But very fast, despite the lack of a decent GPU.

5h agoHN ↗

Prompt for PetitGPT research-v1: "How much wood would a woodchuck chuck if a woodchuck could chuck wood?"

Response: "A woodchuck would chuck wood if it could chuck it. The woodchuck's chucking action is a form of "chucking" or "chucking in" which is a behavior that allows it to extract nutrients from wood. The woodchuck's chucking action is a form of "chucking" because it is a form of "chucking" that allows the woodchuck to extract nutrients from wood."

I chuckled.

4h agoHN ↗

Very interesting but on a naming perspective what is a "micro large" model ?

3h agoHN ↗

Impressive how much you can do locally now. Was expecting much slower inference but it's surprisingly snappy.

2h agoHN ↗

Interesting. Crashed macOS. That's already impressive.

1h agoHN ↗

Sorry. That shouldn't happen. It doesn't use a lot of memory and only uses WebGPU in the normal way. What version of macOS and Safari did you use, and what is your hardware, please?

2h agoHN ↗

Pretty neat how far these small models have come. Running them locally in the browser feels like magic for quick experiments.

1h agoHN ↗

Ran a few of these, and the speed for in-browser execution is wild. Perfect for quick demos without a backend.

42m agoHN ↗

I did a WebGPU-based HTML Single-File Chat Application even for larger models. It just depends on your machine. It is cool to show people what can be done on their machines without installing anything. It also works in air-gapped environments.

https://github.com/moooff/HermitUI