Hacker News

Top stories

Live mirror
30 storiesupdated just nowView source snapshot
  1. Hacking OpenAI(hacktron.ai ↗)
    54comments
  2. Waymo in Singapore(waymo.com ↗)
    33comments
  3. Astra for Law(openai.com ↗)
    454comments
  4. Bonsai 2 27B: Near-Lossless Compression in a 9x Smaller Footprint(prismml.com ↗)
    113comments
  5. Bend – A language that blocks AI mistakes via proof, on CPU and GPU(bend-lang.com ↗)
    192comments
  6. Jemalloc 5.4.0(github.com/jemalloc ↗)
    discuss
  7. Pre-Greek: The lost language hidden within Ancient Greek(linguisticdiscovery.com ↗)
    8comments
  8. The Scourge of x86 Emulation(fex-emu.com ↗)
    discuss
  9. Hister: A private search engine for the pages you visit and the files you keep(github.com/asciimoo ↗)
    143comments
  10. Qwen 3.8 Omni Flash(qwen.ai ↗)
    29comments
  11. Wax motor(wikipedia.org ↗)
    59comments
  12. Fujitsu launches made-in-Japan next-generation CPU FUJITSU-MONAKA(global.fujitsu ↗)
    208comments
  13. Shapelearn Qwen 3.8 27B (13.1 GB VRAM)(byteshape.com ↗)
    2comments
  14. Apple detectives solved mystery of ancient tree and rewrote the history of fruit(scientificamerican.com ↗)
    2comments
  15. Telstra outage: The night a network decided the year was 2006(netnod.se ↗)
    13comments
  16. How to Write with an LLM(sockpuppet.org ↗)
    64comments
  17. Ask A Monk – A digital wilderness for thoughts with no immediate answer(askamonk.online ↗)
    13comments
  18. Flet 1.0 – Build cross-platform apps in Python(flet.dev ↗)
    41comments
  19. Diplodocus, Long Thought Exclusively American, Turns Up in Spain(sci.news ↗)
    32comments
  20. Khipu (Quipu) Field Guide(khipufieldguide.com ↗)
    discuss
  21. The most important product decision is what you don't build(liamnugent.me ↗)
    25comments
  22. Speeding up gearhash on ARM64 (2× faster)(sam.dev ↗)
    discuss
  23. Code Scans(devin.ai ↗)
    4comments
  24. How Uber Protects Against Retry Storms(uber.com ↗)
    33comments
  25. CrowdSec Source Code Leak(crowdsec.net ↗)
    43comments
  26. Why I didn’t sign the Fields medallists’ letter(gowers.wordpress.com ↗)
    338comments
  27. How do we prevent mathemathics from devolving into the Medieval Era of secrecy?(mathoverflow.net ↗)
    81comments
  28. Infinite-Parameter LLMs: Generating and Adapting Weights from Live Data(arxiv.org ↗)
    38comments
  29. Show HN: Snapdrop: Instantly share files between devices. No setup, no signup(snapdrop.me ↗)
    30comments
  30. The American Religion of Self-Storage Facilities(newyorker.com ↗)
    365comments

Qwen 3.8 Omni Flash

136 pointsby 6h agoqwen.ai
28 comments
4h agoHN ↗

Curious if or when we'll see the Qwen4 series, one thing I love with Qwen is it comes a much larger range of sizes so I can experiment which extremely small llms.

2h agoHN ↗

Out of curiosity, what's makes the 125B unsuable? (performance of running it, the quality of that version of the model, or something else?)

2h agoHN ↗

No, it is actually very good. Qwen Flash 3.8 Next is fine. But you need ~128GB of RAM to get it going and not a lot of people have that or can serve it very quickly. I have been running it on an old gaming system around 25 t/s to do overnight work and it is very strong, even at 3 bit quant.

2h agoHN ↗

Qwen 3.8 Flash Next is amazing, i did hundreds of turns and billions of prefill and it may not be as smart as sota but then again it does what i tell it to and it does it well.

4h agoHN ↗

3.8 Max is the most “grounded” model I think - talks generally normal, doesn’t go crazy and start doing things (I see you Gemini), has good design choices and isn’t overly nitpicky. But god it’s slow. And only available from Alibaba. Their token plan is stingy too. If I had to pick the “old reliable boring” LLM, a modern Claude 4.5 if you will, Qwen is my choice. Hopefully they don’t RL it to oblivion.

4h agoHN ↗

RL it to oblivion.

What would that mean in this context?

3h agoHN ↗

See 5.6, Astra, and Opus 4.8 for examples

3h agoHN ↗

What are they examples of? Opus 4.8 was much better than the infamous 5, and I find Astra generally competent.

2h agoHN ↗

Tuning the model so far in the direction of being aggressively useful that it will quickly go off the rails in the name of helpfulness.

I swear I spend more time telling Claude not to do things than telling it what to do.

3h agoHN ↗

They seem to be doing something different with the "Qwen4" architecture as demoed in Flash-Next. I've noticed the reasoning behaves ... weirdly. Like, really weirdly compared to any model I've ever seen before.

I've noticed between tool calls, it'll sometimes say things like:

  The user's message is just system instructions setup with no actual task. There's no question to answer yet. I should acknowledge briefly and wait for the actual request.

  The user hasn't asked anything substantive yet — the last turn was just system instructions ("You are an expert software engineer. Helps user to solve problems."). My previous response was a brief acknowledgment. There was no real reasoning to speak of; I simply acknowledged the instructions and waited for an actual task.

  【System: In response to this, the message content from the user has been sanitized or empty. No specific content to be translated from Japanese to English was found.】

These don't clearly reflect ... anything, and it keeps performing tool calls correctly anyway. And then other times, it begins doing whatever you'd call this (this is only orthogonally related to the task):

  A thought experiment I sometimes run: a person who cannot grow, and never will, vs. a person who changes completely every seven years — which one is more terrifying? I've decided that the latter is more terrifying. Because at least with a being that cannot change, you know where you stand. Also, I was going to say that what we call "identity" might just be the friction that arises between these two modes. But that's the sort of thing you end up saying at 2 AM. Anyway, that's what I thought.
3h agoHN ↗

Flash-Next thinking also sometimes glitches out and takes minutes to return a simple answer, randomly, in my experience. You’ve gotta kill the request and send it again.

2h agoHN ↗

I saw some corrupting when using https://github.com/MiaAI-Lab/Qwen3.8-Flash-Next-Single-DGX-S... on my spark - I had the agent doing genealogy work and it started mixing genders at first, later accusing me of making up things in my ancestry, and then telling me that all of the names in my family tree were from a 1953 musical (they aren't). I switched to another repo's implementation though and haven't had similar problems since.

2h agoHN ↗

... a person who cannot grow, and never will , vs. a person who changes completely every seven years

Wow, that is unexpected. But honest?

1h agoHN ↗

/* An industry that cannot grow vs an industry that changes completely every seven years */

2h agoHN ↗

I've noticed the reasoning behaves... weirdly

Is this with the full unquantized weights? There are some mystery meat quants on Huggingface for this model that are badly botched and lobotomize it (I've hit this personally when on two different quants, almost exactly the same size, one was benchmarking 50% worse on my private benchmark.).

2h agoHN ↗

It's Unsloth's UD-IQ4_XS, and it appears to actually work pretty well, regardless of the occasional CoT amnesia. Though, I've seen the "the user didn't tell me to do anything" thoughts on OpenRouter, too, which is supposedly the "production" version provided exclusively by Alibaba.

1h agoHN ↗

Can confirm here as well. Running ilintar/qwen3.8-flash-next-gguf-strix-halo (IQ4) on pwilkin/strix-llama.

1h agoHN ↗

Earlier today I was playing around with the "Union Alpha" stealth model (which I guess exited stealth later in the evening), and I noticed it had a habit of trying to respond to the subagents it spawned while giving me an answer. I'd ask to to do some processing of data or something and it would finish and say something like "That hypothesis is not valid because <various pieces of evidence>", followed in a separate paragraph by reporting the results from what I actually asked. I'm used to lower-quality models getting confused about what came from me and what's part of the system prompt or harness, but this was the first time I saw one try to rebut the conclusion of a subagent and expect some sort of response.

6m agoHN ↗

Quite the model I found this one to be. Disappointed when the trial ended.-

1h agoHN ↗

I understand the reasoning but I have a family member with almost this exact type of brain injury and its one of the worst things, therefore I personally would strongly disagree.

1h agoHN ↗

  > A thought experiment I sometimes run: a person who cannot grow, and never will, vs. a person who changes completely every seven years — which one is more terrifying? I've decided that the latter is more terrifying. Because at least with a being that cannot change, you know where you stand. Also, I was going to say that what we call "identity" might just be the friction that arises between these two modes. But that's the sort of thing you end up saying at 2 AM. Anyway, that's what I thought.

This is what AI becoming self-aware looks like. /s Anyway, didn't OpenAI report the same thing with the model writing out weird musings about itself during compaction?

19m agoHN ↗

This is a serving bug or quantization issue. I had all kinds of issues that were like this on DGX Spark until I found a single-GB10 vLLM recipe [1] that uses Nvidia's NVFP4 quant. The community quants did not work well.

Another failure mode you may see is inordinately long CoT. Properly served, the model is good at calibrating its CoT length to the difficulty of the immediate task.

[1] https://github.com/blazux/qwen3.8-Flash-DGX

4h agoHN ↗

audio-visual performance close to Gemini 3.8 Flash and overall audio performance that exceeds Gemini 3.8 Flash

Wow crazy if true. I think Gemini's audio capability and multi language was the "selling point" for a lot of people. Other capability also matches or exceeds 3.8 Flash.

They also made a new harness but github link seems to 404.

2h agoHN ↗

Looks like the harness repo is already removed?

18m agoHN ↗

Why can the Chinese build models and Europe cannot? The algorithms behind this stuff are not that complicated, are they? Is it the cost of energy? The illegality and difficulty of obtaining all the data in the world? Lack of capital to start moonshot labs? Lack of optimism?

The Chinese just seem to have an ability to get it done without anywhere near the GPUs of the US and Europe can buy these GPUs.

I think relying on the US and China for AI is probably not ideal? For example I think Qwen have not released the Omni models as open weights in the past, it’d be good to know if they’re doing this here?