Hacker News

New stories

Live mirror
30 storiesupdated just nowView source snapshot
  1. Show HN: Advance Sleep Cycle Calculator(globaltoolsbox.online ↗)
    discuss
  2. Googlebook – Meet the Lineup and pre-order(googlebook.google ↗)
    discuss
  3. Grok 4.7 Intelligence, Performance and Price Analysis(artificialanalysis.ai ↗)
    discuss
  4. Show HN: Agent Chaperone – Screen AI agent tool calls and results with Jev(github.com/agent-chaperone ↗)
    discuss
  5. Show HN: flyOS – A fruit fly connectome simulated in real time on an iPhone(becomethefly.com ↗)
    discuss
  6. Saudi Arabia's Ceer launches flagship electric vehicles(agbi.com ↗)
    discuss
  7. Modulate ML Team Announces New Public Entity Transcription Benchmark(modulate.ai ↗)
    discuss
  8. WTF Is Up with Napster's AI Pivot?(tedium.co ↗)
    discuss
  9. Alcor: Simulate cpuid and sgdt/sidt results per-process(github.com/er-azh ↗)
    discuss
  10. Google hit with €403M fine by Irish data watchdog over GDPR violations(bbc.co.uk ↗)
    discuss
  11. Raspberry Pi founder Eben Upton: 'I'm an Omni-geek(ft.com ↗)
    discuss
  12. Grok 4.7 is here with Electrical engineering benchmark which beats fable 5.1 max(twitter.com/hive_echo ↗)
    discuss
  13. Delta A21N at Kahului on Sep 19th 2026, fuel fumes on board(avherald.com ↗)
    discuss
  14. WWLD #1: Domains and gas station hotdogs(chaosguru.substack.com ↗)
    discuss
  15. The agents, they just want to talk(snats.xyz ↗)
    1comments
  16. Amazon Blocks Meta's Muse AI Agent from Its Retail Site(bloomberg.com ↗)
    1comments
  17. Show HN: Foremerge – Catch Intent Conflicts Between Parallel Coding Agents(github.com/naw103 ↗)
    discuss
  18. Measuring nothing (with great accuracy) (2014)(seths.blog ↗)
    discuss
  19. Thinking, Fast and Slow(wikipedia.org ↗)
    discuss
  20. Building a Browser Without V8: What Broke, and What Worked(medium.com/koukyosyumei ↗)
    discuss
  21. Show HN: Fusion-runtime – self-hosted voice agents, STT+LLM+TTS in one process(github.com/samarthurs18 ↗)
    discuss
  22. Show HN: Code Graph View – What if you could navigate a codebase as a graph?(github.com/otobongfp ↗)
    discuss
  23. How Secrets Work in Docker(infisical.com ↗)
    discuss
  24. Steerable Cultural Preference Optimization of Reward Models(arxiv.org ↗)
    discuss
  25. pwn.college - Learn to Hack(pwn.college ↗)
    1comments
  26. First Autonomous Flight Across the United States(jobyaviation.com ↗)
    1comments
  27. This Digital Radio Gets Messages to the World’s Remotest Locations(ieee.org ↗)
    discuss
  28. Fable 5 – Median thinking declined in August(twitter.com/lon ↗)
    discuss
  29. Nginx Control API: View In-Memory Configuration and Reload via HTTP Requests(nginx.org ↗)
    1comments
  30. The reason children no longer enjoy reading(spectator.com ↗)
    discuss

Grok 4.7

58 pointsby 49m agox.ai
21 comments
37m agoHN ↗

after using cursor grok & trae.ai for several months , grok curor is highly superior results to trae.ai

28m agoHN ↗

I guess it’s only my opinion but having used grok for personal chat: it’s by far the worst one amongst Claude, ChatGPT and even Deepseek, Gemini etc.

The personality is bland and it doesn’t work nearly as hard or even tries to help.

16m agoHN ↗

I used openrouter to send same prompt to qwen, derpseek, gemini and grok and found that grok does good research and produces less bullshit, especially when prompted to be critical of an idea

15m agoHN ↗

This has been my experience as well. Grok will end tasks almost immediately and claim "Done!". It's definitely the laziest and most "dishonest" of all the models. The others aren't perfect, but I can't use Grok for any serious coding task.

11m agoHN ↗

The personality is bland

I don't use Grok, but do you want your LLM to have a personality? "Personality" is exactly what people don't like about Claude.

24m agoHN ↗

Apparently Grok 4.7 has 40% more weights than Grok 4.6, but the price ($6 output token, $2 input) is the same.

Given that the decrease in their margin and the fact they delayed the release of Grok 4.7 almost two weeks past the original date, XAI must not have been happy with the results for 4.7. And XAI also waited the day before Opus 5.5 is rumored to launch. I imagine Opus 5.5 will blow Grok 4.7 out of the water benchmark wise.

However, I have become skeptical of benchmarks. Grok 4.5 solved some issues setting up a buildroot system that Fable 5 couldn't do. I find the post cursor groks are phenomenal at frontend web development, though Claude is much better at backend ruby.

My favorite part of the new Groks has been how they speak in plain english. I simply cannot stand Claudish. Or even GPT, which doesn't have Claude's ticks but definitely likes to handwave explaining technical concepts. Still, nothing beats Claude 3.5 and 4 with explaining since it seems all models have regressed. I wonder if Grok 4.7 will also regress with English because of all the RL.

14m agoHN ↗

For what it's worth - over the last few years or whatever, it seems like Anthropic benchmaxxes the least.

That being said, I currently prefer Sol / Astra to Opus / Fable as I find both to be a better cost payoff to me.

6m agoHN ↗

Token price doesn't tell you much without knowing token efficiency.

21m agoHN ↗

What's with the deceptive graph on top? Not including Astra can't have been an oversight, did the model compare poorly to it?

17m agoHN ↗

I wonder if that means that SpaceX evals show that they consider astra better than fable or that they hate Sam&co so much they don't want to show their stuff.

13m agoHN ↗

They have Astra in other benchmarks lower on the page. They just don't want to show it winning

11m agoHN ↗

The chart is cursorbench though and they asked about the "deceptive graph"

7m agoHN ↗

Deceptive? An extremely quick google search would answer your question. OpenAI pulled out of Cursor before they released Astra so it never got that benchmark.

16m agoHN ↗

How anyone that values democracy in the United States could support any of Elon's ventures is difficult to understand.

7m agoHN ↗

Judging by Elon's staggering success in all of his ventures I'd say you're out of touch.

5m agoHN ↗

The public very much voted for massive administrative reform. Are you refering to DOGE, Elon Musk's influence on elections, something else?