Hacker News

New stories

Live mirror
30 storiesupdated just nowView source snapshot
  1. Reverse Engineering How Meta's Muse Shops (caeliai.com)
    —discuss
  2. Anthropic's IPO prospectus shows AI vision, surging costs (reuters.com)
    —discuss
  3. Nvidia announces AI safety platform (theverge.com)
    1comments
  4. Holo4: Powering generalist computer-use agents (huggingface.co)
    —discuss
  5. Pklm-sandbox – A lightweight open-source LogitsProcessor for local LLMs (github.com/theimmortalpython)
    —discuss
  6. Conrad Barski wrote a quirky file explorer (twitter.com/lisperati)
    —discuss
  7. EVs being trialled as giant batteries to power homes for less in Auckland (rnz.co.nz)
    —discuss
  8. Come Try the Srena (thearena.rip)
    —discuss
  9. Mark (colossus.com)
    —discuss
  10. Modern Warfare 2, Skate 3 and Minecraft in one game, on IW4L, MW2 Rust rewrite (github.com/chasmlol)
    —discuss
  11. WSJ reports OpenAI scrapped GPT-6.1 Astra over safety concerns (runtimewire.com)
    —discuss
  12. US finalizes new lower fuel economy standards (reuters.com)
    —discuss
  13. Google seemingly confirms plans to kill ChromeOS in 2034 (arstechnica.com)
    —discuss
  14. Agent token spend distribution: coding versus legal work (lexifina.com)
    —discuss
  15. Show HN: Tiny e-ink display for your portfolio (snaptra.de)
    —discuss
  16. Show HN: MuseTogether – let Meta's Muse agents coordinate plans with each other (musetogether.ai)
    —discuss
  17. MoE Analysis Qwen[3.5|3.6]35B-A3B (gslaller.github.io)
    1comments
  18. When we talk about traffic violence (maxmautner.com)
    —discuss
  19. Check twice, cut once with LLM search relevance eval (softwaredoug.com)
    —discuss
  20. AI tools generated nearly $1B in extra costs, Blue Cross insurers say (reuters.com)
    1comments
  21. Block US sanctions on ICC, EU lawmakers tell Commission (politico.eu)
    —discuss
  22. Training NanoGPT in 39.9 Seconds (hyperstition.cc)
    1comments
  23. Show HN: LLMs Play Chicken in Realtime (wildcardlabs.tech)
    —discuss
  24. Sonnet 5.5 scores just behind Opus 5.5 on Artificial Analysis Intelligence Index (artificialanalysis.ai)
    —discuss
  25. The Ratchet (rubick.com)
    —discuss
  26. Tech OpenAI abandons plan to release upcoming model as safety concerns escalate (cnbc.com)
    —discuss
  27. Joseph Henrich's Report on Anthropology (vanderbilt.edu)
    —discuss
  28. Mosquitoes Are a Choice (worksinprogress.co)
    2comments
  29. Go Pace Yourself (an example of what AI films done right looks like) (reddit.com)
    1comments
  30. Electronics in U.S. homes by income [pdf] (eia.gov)
    —discuss

MicroLLM Lab – Try 7 tiny LLM's in the browser

102 pointsby 4h agostateofutopia.com
47 comments
4h agoHN ↗

You can try 7 different tiny LLM's in your browser.

3h agoHN ↗

What I missed from the title was this: Can I try 7 different tiny LLMs in my browser?

3h agoHN ↗

i have a bromine tub, give me simple instructions for what to do since i just filled it up w fresh water

1. Add 1 tablespoon of water to the water bath. 2. Place the tub into the water bath and let it sit for about 5 minutes. 3. After 5 minutes, remove the tub and let it cool down. 4. Now, fill the tub with water and let it sit for about

uhmm completely unusable ?

3h agoHN ↗

No, that is not what SLMs are useful for. They have very little knowledge, and typically lack reasoning skills.

Useful applications include sentiment analysis, text classification, entity extraction, etc.

They certainly can be useful, but you shouldn't compare them with LLMs such as Opus or Fable.

3h agoHN ↗

Completely unusable for that, yeah, that's a 100M model, not an assistant

3h agoHN ↗

The website literally says they are SLM's .. derp af

3h agoHN ↗

I did the default arithmetic with PetitGPT research-v1

What is 2+2?

Answer

To find 2 + 2, we need to add 2 to both sides of the equation.

2 + 2 = 4

So, 2 + 2 = 4 + 2.

Brilliant

2h agoHN ↗

That's not incorrect.

LLMs produce semantically correct sentences, not factually correct statements. Have we forgotten this so soon?

2h agoHN ↗

It is, in every sense, incorrect. Which statement in this short snippet is "semantically correct"? (Better LLMs get this right, of course.)

1h agoHN ↗

It's not very good semantically either.

Give me a recipe for soup.

Here is a recipe for soup:

Saffa-Cake-Sweet-Sweet-Sweet-Sweet-Sweet-Sweet-Sweet-Sweet-Sweet-Sweet-Sweet-Sweet-Sweet-Sweet-Sweet-Sweet-

2h agoHN ↗

GPT-2 124M is terrible...

  > what is 2+2?

Answer:

  > 3+3? 4+4? 5+6? 7+8?
  > Reply ~18000 0 ~10 min 2 By : 1-1: I'm a beginner. 3x2 is my best option, but if you're not sure about the other options then just go for it and try again
2h agoHN ↗

It's not instruct tuned looks like? It's closer to a base model rather than a chatbot.

1h agoHN ↗

That one is a 2019 model :) Years before the ChatGPT public preview.

1h agoHN ↗

GPT-2 is an interesting one because it is a February 2019 model. (You can see some information about it below the card if you click on the card.)

That was 2-3 years before the big "ChatGPT moment" (the highly coherent ChatGPT research preview was released in November 2022, I think it was ChatGPT 3.5). Back in 2019 the models really were not producing very coherent output. Now you can see it for yourself right in your browser :) Everything has come a really long way since then!

1h agoHN ↗

I've been playing with it for a bit, and I'm a tiny little bit surprised that they decided to continue pursuing that direction of research at all. What I'm getting from it looks like it could just be arbitrary sentence and paragraph fragments from the Internet, pasted together Markov-chain like.

I'm not sure I would have ever believed that something useful would come out of it, yet here we are.

1h agoHN ↗

Awesome! I tried it and after loading (which took a while as it is a large model) got 12 tokens/second and very coherent output. Great demonstration.

40m agoHN ↗

I would recommend adding the MiniCPM5-1B model. Surprisingly coherent for a 1B model and performs well. On my pixel 9 I get 33 tok/s on CPU. On GPU I get about 26 tok/s but prefill jumps to nearly 500.

3h agoHN ↗

Cool project, but I'd really suggest looking at the UI.

The text is too small and it's way too dense with information in general. Considering how simple this product is to use, it's kinda crazy that I have to scroll through over a page length of (mostly useless, AI-generated) information before getting to the actual interface.

Also what is going on with the footer (why does it link back to the site itself, why is it telling me to "serve over HTTP").

2h agoHN ↗

This is the future of software, sloppy ui.

2h agoHN ↗

Thank you for the feedback, I'll think about how to incorporate the changes you've suggested.

Edit: I've concluded that the text is tiny so people can skip it, it's like fine print. There are only 600 words on the entire page, that's not a lot even if someone does read all of them.

Since this submission has been on the front page of HN for hours now, clearly people like it as it is. I don't think I need to change anything about it.

The footer means that if you download the zip file and unzip it, you can't just open index.html and expect it to work, you need to start a local webserver for it.

26m agoHN ↗

It's clear that this was vibecoded. I'm antivibecoding, but here's a tip.

Don't focus on the actual code or the actual UI. You need to give your model a way to see its output. If it just pumps code, it looks at the code. If you let it look at the rendered UI, whether by manually taking screenshots, or making it so that the agent looks at the browser output automagically (with a webdriver or selenium or puppeteer or a skill or whatever), it will fix whatever issues it sees on itself, and it will never be content with delivering something that looks as bad as this.

3h agoHN ↗

PetitGPT told me that

"2+2 is 2."

Otherwise, a very neat demo. As others have said, the UI is VERY confusing, way too much stuff going on.

3h agoHN ↗

Really cool project. Giving web apps direct access to on-device models is something I’m excited about, and it’s cool to see the different approaches.

I’ve been working on a related proposal called the Web Models API, which explores a browser standard for an API that runs open-weight models on-device. Would love your thoughts: https://www.webmodels.dev

2h agoHN ↗

I read your proposal, I think it's great! Where will the navigator get the model if the user agrees to download it? For this demonstration I just serve the models on my own server, but for larger models it may be an issue as they may not have direct download links even if they are open weights.

2h agoHN ↗

Great question. Right now there isn’t a definitive answer, but it’s something that needs to be worked out. There would likely be a registry. The question is how to keep model IDs consistent: does each browser manage its own registry, or is there one shared across browsers?

2h agoHN ↗

Since you're asking for some kinds of permissions anyway, you could ask if the user is willing to also seed the model, p2p. (However, seeding files is not as popular as it used to be, many residential Internet connections don't have good upload.) If you have the capacity for it, your site webmodels.dev could act as a tracker and initial seed for any models. Then it could be the one central registry. It might get to be too much for you though, a lot of the open weights models are huge.

2h agoHN ↗

Interesting idea. I hadn't considered that. Thanks

2h agoHN ↗

unfortunately on firefox: Uncaught ReferenceError: GPUShaderStage is not defined

1h agoHN ↗

I tested it on Firefox on windows, version 156.0.1 and didn't get that error.

What version of Firefox are you using and what is your operating system and graphics card, please? Can you also try it without WebGPU? (Reload the page and uncheck "Prefer WebGPU" and try a prompt.)

7m agoHN ↗

Same issue. Latest version of Firefox 156.0 on Ubuntu GNOME. AMD Radeon 860M.

Unticking the WebGPU checkbox and reloading didn't do anything.

1h agoHN ↗

It's a cool project, but I couldn't get it to load. What did you test this on? I tried the live link here:

https://willaaam.github.io/gemma-4-E2B-webgpu-vision/

And after loading it, with an NVidia 1060 GPU (6 GB RAM) on Windows it failed with "Failed to load: No supported WebGPU variant for com.xenova.gemma4.DenseGemv; rejected sgma".

In Safari on a 2026 Mac Mini M4 with 24 GB of RAM it failed with "Failed to load: JSON Parse error: Unexpected EOF".

The idea is pretty cool though!

25m agoHN ↗

Huh, weird! It doesn't work in Firefox, but I test on Safari (M5 Max) and Chrome (Linux, Arc B580). I don't have access to NVidia or AMD hardware myself unfortunately, though friends confirmed both as working.

Going to debug tomorrow!

1h agoHN ↗

What browser are you using? I tested it on Windows, Mac, and iPhone. I tested it in Chrome, Firefox, Edge, and Safari. Everything works on the three machines and phone I tested it on.

(It's a little bit slow at the moment - you have to wait a few seconds for the models to load - as it's currently on the HN front page. The server is on a 1 gigabit unmetered network connection so it can serve all the weights - around 600 megabytes - to one person every few seconds, there are several concurrent users now.)

1h agoHN ↗

how would you compare your capabilities to that of claude fable 5.1 by anthropic

Comparing your capabilities to that of claude fable 5.1 would be very similar. Both are stories about a clown who is a cunning and manipulative animal. They both have problems and challenges, and they face challenges to overcome them. However, they have different goals and approaches. The clown

I think that tells you all you need to know.

1h agoHN ↗

Thanks. Unfortunately it's too late to change the title!

8m agoHN ↗

Maybe we should just call them SLMs instead of LLMs...