- 380comments
- 82comments
- 36comments
- 167comments
- 139comments
- 50comments
- 198comments
- 7comments
- 3comments
- 36comments
- 22comments
- 1comments
- 2comments
- 41comments
- 129comments
- 19comments
- 25comments
- 35comments
- 316comments
- 107comments
- 60comments
- 181comments
- 347comments
- 36comments
- 5comments
- —discuss
- 20comments
- 4comments
- 51comments
- 104comments
Curious if or when we'll see the Qwen4 series, one thing I love with Qwen is it comes a much larger range of sizes so I can experiment which extremely small llms.
3.8 Max is the most “grounded” model I think - talks generally normal, doesn’t go crazy and start doing things (I see you Gemini), has good design choices and isn’t overly nitpicky. But god it’s slow. And only available from Alibaba. Their token plan is stingy too. If I had to pick the “old reliable boring” LLM, a modern Claude 4.5 if you will, Qwen is my choice. Hopefully they don’t RL it to oblivion.
What would that mean in this context?
See 5.6, Astra, and Opus 4.8 for examples
What are they examples of? Opus 4.8 was much better than the infamous 5, and I find Astra generally competent.
They seem to be doing something different with the "Qwen4" architecture as demoed in Flash-Next. I've noticed the reasoning behaves ... weirdly. Like, really weirdly compared to any model I've ever seen before.
I've noticed between tool calls, it'll sometimes say things like:
These don't clearly reflect ... anything, and it keeps performing tool calls correctly anyway. And then other times, it begins doing whatever you'd call this (this is only orthogonally related to the task):
Wow crazy if true. I think Gemini's audio capability and multi language was the "selling point" for a lot of people. Other capability also matches or exceeds 3.8 Flash.
They also made a new harness but github link seems to 404.