Hacker News

New stories

Live mirror
30 storiesupdated just nowView source snapshot
  1. Dreams of making drugs and tissues in space get closer to reality(science.org ↗)
    discuss
  2. Vulnerability Disclosure Policy(flocksafety.com ↗)
    discuss
  3. Tinfield 1 is an open weight coding model from Nigeria that beats Opus 4.8(twitter.com/badtheorylabs ↗)
    discuss
  4. 2026: The Year Galleries Realised the Need to Flag or Cull Slop Images(techrights.org ↗)
    discuss
  5. Big AI to humanity: drop dead(matthewbutterick.com ↗)
    discuss
  6. Ten years since the factors (fiction)(dreamstation.systems ↗)
    discuss
  7. The Apocalypse Will Not Be Sexy(techdirt.com ↗)
    discuss
  8. Qwen3.8-Flash-Next on a 64 GB M2 Ultra: A 66-Minute Real Work Run(b1tank.github.io ↗)
    discuss
  9. Building standards for the next phase of AI(openai.com ↗)
    discuss
  10. Seeing Who Contaminates Linux with Slop (and Also Admits It)(techrights.org ↗)
    discuss
  11. Skyportal CLI: What happens when an AI agent needs approval(pypi.org ↗)
    discuss
  12. Show HN: Sommelier, an Excel Addin for your agent(grokked.it ↗)
    discuss
  13. Seven Deadly Sins of DevX(amplitude.com ↗)
    discuss
  14. Gobag: Semantic session archival for Claude Code(github.com/satmihir ↗)
    discuss
  15. Bernstein's Factorization Method Helped Factor RSA-240 in 2020(leetarxiv.substack.com ↗)
    discuss
  16. Petal: Building the First Petabit-Class Transoceanic Subsea Cable(fb.com ↗)
    discuss
  17. OpenAI projections point to a $278B cash burn through 2030(tomshardware.com ↗)
    1comments
  18. LGBT tolerance slumps in the Netherlands, 22% of boys have LGBT+ positive views(dutchnews.nl ↗)
    discuss
  19. Aura – Open-Source Framework for Detecting Social Engineering in LLMs(github.com/kate8382 ↗)
    discuss
  20. What would happen if the Yellowstone supervolcano erupted now(theconversation.com ↗)
    discuss
  21. How do I talk about using AI without sounding like an AI evangelist?(askamanager.org ↗)
    discuss
  22. Writing Rust code faster than SotA by asking agents to make the code faster(minimaxir.com ↗)
    discuss
  23. How to address the failures we found along the US border's "virtual wall"(technologyreview.com ↗)
    discuss
  24. France to Upgrade Heat Wave Modeling After Record Nuclear Curbs(bloomberg.com ↗)
    1comments
  25. AI proves Medvedev logic is undecidable(arxiv.org ↗)
    discuss
  26. Amazon Blocks Meta's New Muse AI Agent from Shopping on Amazon.com(forbes.com/sites/jonmarkman ↗)
    discuss
  27. I don't care about majors, minors, or honors programs. Do you?(computationalcomplexity.org ↗)
    discuss
  28. Ask HN: How close are we to AI bubble bursting?
    1comments
  29. The Questionable Legality of ICE's Palantir Elite System(techpolicy.press ↗)
    discuss
  30. The Influence of AI on Human Decisions in DHS Surveillance(techpolicy.press ↗)
    discuss

Grok 4.7

181 pointsby 1h agox.ai
90 comments
1h agoHN ↗

after using cursor grok & trae.ai for several months , grok curor is highly superior results to trae.ai

1h agoHN ↗

I guess it’s only my opinion but having used grok for personal chat: it’s by far the worst one amongst Claude, ChatGPT and even Deepseek, Gemini etc.

The personality is bland and it doesn’t work nearly as hard or even tries to help.

1h agoHN ↗

I used openrouter to send same prompt to qwen, derpseek, gemini and grok and found that grok does good research and produces less bullshit, especially when prompted to be critical of an idea

28m agoHN ↗

Ask it to be critical of the birthday photos and see where that gets you.

1h agoHN ↗

This has been my experience as well. Grok will end tasks almost immediately and claim "Done!". It's definitely the laziest and most "dishonest" of all the models. The others aren't perfect, but I can't use Grok for any serious coding task.

1h agoHN ↗

The personality is bland

I don't use Grok, but do you want your LLM to have a personality? "Personality" is exactly what people don't like about Claude.

1h agoHN ↗

Apparently Grok 4.7 has 40% more weights than Grok 4.6, but the price ($6 output token, $2 input) is the same.

Given that the decrease in their margin and the fact they delayed the release of Grok 4.7 almost two weeks past the original date, XAI must not have been happy with the results for 4.7. And XAI also waited the day before Opus 5.5 is rumored to launch. I imagine Opus 5.5 will blow Grok 4.7 out of the water benchmark wise.

However, I have become skeptical of benchmarks. Grok 4.5 solved some issues setting up a buildroot system that Fable 5 couldn't do. I find the post cursor groks are phenomenal at frontend web development, though Claude is much better at backend ruby.

My favorite part of the new Groks has been how they speak in plain english. I simply cannot stand Claudish. Or even GPT, which doesn't have Claude's ticks but definitely likes to handwave explaining technical concepts. Still, nothing beats Claude 3.5 and 4 with explaining since it seems all models have regressed. I wonder if Grok 4.7 will also regress with English because of all the RL.

1h agoHN ↗

For what it's worth - over the last few years or whatever, it seems like Anthropic benchmaxxes the least.

That being said, I currently prefer Sol / Astra to Opus / Fable as I find both to be a better cost payoff to me.

39m agoHN ↗

I was going to say the reverse - claude has been the less satisfying normalized by benchmark for me in the last year. Both astra and fable have their quirks, but I am 90% codex this year up from 10% last year.

55m agoHN ↗

Token price doesn't tell you much without knowing token efficiency.

36m agoHN ↗

Their leading benchmark with cost per task shows a tough sell compared to Fable 5.1 Low and doesn't reach the performance of Fable 5.1 Medium.

How representative that is of real world usage, I don't know.

In their benchmark GPT 5.6 Sol performs suspiciously poorly compared to the former models.

36m agoHN ↗

I simply cannot stand Claudish

I totally agree, it’s like that as models become more intelligent, they are less understandable by most of people... but aren’t we humans doing the same?

33m agoHN ↗

The best ideas are usually the simplest to elaborate. If someone comes up with a convoluted scheme that are hard to understand or be adequately explained, it's usually fraud.

When claude speak in convoluted mess, they are often going off on tangents in real work that you asked it to do, too.

24m agoHN ↗

That believes that the world can be simplified into dichotomies, or at least, simplified. Sometimes problems are complex, and the solutions to them necessarily so. For example, cancer. I order to begin to understand that problem, you have to understand the utter complex scheme it has devised in order to exist. A 20 minute YouTube video isn't going to be able to begin to cover the basics of the subject, although there are some good ones, with clever analogies.

Just because something is difficult to understand doesn't mean it's fraud, although if someone is trying to dazzle you with clever words and names of institutions you recognize because they are selling you something, there's a good chance they're lying to you in order to get some money from you.

7m agoHN ↗

No but almost all good ideas can be reduced down to a few sentences if you're good at explaining things. It's a different kind of intelligence than what's commonly called IQ but it's something like that regardless.

Sure the explanation will oversimplify a lot but then you can expand it recursively if needed, you gotta start somewhere.

28m agoHN ↗

If you can't explain it simply, you don't understand it well enough

23m agoHN ↗

That's half true. A very smart model should be able make good explanations, which include simple understandable prose. That can should be possible even as its thought process gets more alien.

16m agoHN ↗

Agreed. The more knowledge you amass on a subject, the more important it becomes to be extremely specific and nuanced - or your communications end up being incorrect. You become better at expressing your thoughts, but harder to understand.

The weird thing is, that's not what AI models seem to be doing. The prose is just weird.

8m agoHN ↗

You become better at expressing your thoughts, but harder to understand.

This happens most though when the speaker doesn't (or care to) understand their audience.

Eg i find effective communication requires expertise in both the subject matter domain but also the reference of the listener. Eg in ELI5 framing, if you don't know what information 5yr olds are expected to know you'll do a poor job at an ELI5.

It often feels like Claude does poorly at both framing the response relative to what it "thinks" the listener knows, but also the prose is... sideways, just weird as you said.

15m agoHN ↗

What I notice about Claudish is that it has its preferred cliche’s and overstretched methaphores, it packs too many ideas in a sentence, and to achieve the latter it makes up adjectives.

I should try adding these tips to my system prompt. Is there a shorthand to describe such language use? I am not a native English speaker.

35m agoHN ↗

I expect the next Anthropic release to finally reduce the prevalence of Claudish

11m agoHN ↗

It's pretty much the biggest complaint of Claude compared to its competitors, so they really should adress it .

28m agoHN ↗

If they fix Claudish, they've earned me back as a max customer!

Fable 5.1 is not there quite there yet.

They need to get that Sonnet 3.5 magic back.

17m agoHN ↗

Same. The issue with Anthropics models is that (speaking regarding code generation) they REFUSE any kind of comment override instructions. I've tried everything and no matter what, after a few turns, they resort to generating the same overtly verbose junk. Bun's codebase is littered with them See

   // `HANDLE` is an opaque kernel handle (kernel32 validates and returns 0/FALSE
   // on a non-console handle); every out-param is `&mut T` to a `#[repr(C)]` POD,
   // ABI-identical to the Win32 `LP*` pointer (thin non-null). The reference type
   // encodes the only pointer-validity precondition, so `safe fn` discharges the
   // link-time proof. (`bun_windows_sys::kernel32` declares these with `*mut`;
   // redeclared locally so the legacy-conhost cursor path below is plain calls.)

or

   // Progress's terminal handle is the canonical `output::File` (vtable-backed
   // stderr/File from `OutputSinkVTable`). The duplicate `ProgressTerminalVTable`
   // from B-0 round 1 is removed; tty/ansi/winsize route through    the new
   // `OutputSinkVTable` slots so `bun_core` stays T0 (no `bun_sys` dep).

from src/bun_core/Progress.rs

18m agoHN ↗

Perhaps you haven't had the chance to use it, but 3.8 flash is the best model for talking too. Even routing Claudes output through 3.8 to have it explain whats going on is a breath of fresh air

17m agoHN ↗

it's definitely not bigger. smaller if anything looking at how much faster it is

1h agoHN ↗

What's with the deceptive graph on top? Not including Astra can't have been an oversight, did the model compare poorly to it?

1h agoHN ↗

I wonder if that means that SpaceX evals show that they consider astra better than fable or that they hate Sam&co so much they don't want to show their stuff.

1h agoHN ↗

They have Astra in other benchmarks lower on the page. They just don't want to show it winning

1h agoHN ↗

The chart is cursorbench though and they asked about the "deceptive graph"

17m agoHN ↗

They can benchmark it because you can use an openai api key with cursor. Astra is just not included in the cursor plan.

41m agoHN ↗

That's not accurate. OpenAI doesn't allow Grok to provide Astra to Cursor customers anymore, but it doesn't ban anyone from using Astra via alternative harnesses.

If Cursor wanted to include Astra in CursorBench nothing would stop them, they could easily have spent half an hour vibecoding in OpenAI API key support - if it hadn't been convenient to neglect to do that.

37m agoHN ↗

Even if they could do that (workaround to include Astra in CursorBench), that has no practical consequences for Cursor users and that's what I as a Cursor user (what I use for dev, though I use ChatGPT for non-dev stuff) care about.

36m agoHN ↗

It would make the benchmark way better obviously, by showing how their new model compares to their competitors, the whole point of benchmarks and graphs.

12m agoHN ↗

The point of Cursor Bench is to show how models perform in Cursor. If 99% of their users won't be able to access a model unless they go out of their way to include setup an API key for it (which would be insanely expensive with Astra), why would they include it in the benchmark?

29m agoHN ↗

with a proposed shutoff date of November 12, 2026

That said, I don't expect them to benchmark Astra in their Cursor harness given the situation.

10m agoHN ↗

Cursor never added Astra to its consumer subscription plans. And it's likely exactly because of this announcement. Why would they add support for a model they would have to remove shortly after?

56m agoHN ↗

Deceptive? An extremely quick google search would answer your question. OpenAI pulled out of Cursor before they released Astra so it never got that benchmark.

38m agoHN ↗

Pulled out from letting them resell Astra access, that's not a limitation on running a benchmark.

51m agoHN ↗

In my experience Grok especially inside Grok build is pretty solid choice, it’s a no nonsense model and stays on its course. Another surface where I truly enjoy the experience of using Grok model is Grok bot

6m agoHN ↗

I've had really good experiences with Grok 4.6 and grok build. I've been playing around with tscircuit and it can write code with an understanding of spacial reasoning, while also importing cad components from different file formats into tsx, I've been having claude come in and try to error check it and so far claude hasn't found anything to improve in my three projects.

I'm excited for 4.7 although I share skepticism with other users whether 4.7 will be significantly better, since they didn't raise the price.

41m agoHN ↗

Nice to see this release cadence increasing and some continued improvement in quality. I am guessing these models are basically still outcomes of the cursor team integrating with the massive amount of compute they now own: I’d imagine we will see significant step up improvements with grok 5 later this year as the team gets more experienced and confident with larger training deployments. Here’s hoping for another competitive frontier model!

38m agoHN ↗

Waiting on simonw "Generate an SVG of a pelican riding a bicycle " benchmark to judge this model

33m agoHN ↗

You mean a performative pseudo-benchmark that tests for nothing.

At this point you might as well ask an AI model to generate audio waveforms from text and judge it as an audio model or ask a model specifically designed to generate SVGs [0] to generate videos.

[0] https://quiver.ai/

12m agoHN ↗

Thank you. If this whole thing isn't fun, it isn't worth doing.

22m agoHN ↗

You don't think it's useful to learn whether a model's "intelligence" generalizes beyond the tasks and modalities it is usually optimized for?

19m agoHN ↗

Absolutely not. Makes about as much sense as judging a car based on how good an airplane it makes.

10m agoHN ↗

Have you considered that the single most impressive breakthrough of LLMs as a technology is their ability to generalize beyond what they were explicitly trained on? Great analogy, pal, but LLMs aren't cars.

9m agoHN ↗

I disagree. If GPT-7 can draw the Mona Lisa in MS Paint via computer use, this would be interesting.

That it isn't the most efficient way to achieve the same end result is irrelevant.

38m agoHN ↗

I tried in Omp (Oh-my-pi), and so far it's really problematic.

It will loop in thinking mode ("Let me implement those fixes: Fix 1, Fix 2, Fix 3 .... Fix 80, Fix 81"), ignore the AGENTS.md instructions, corrupt plan files, etc etc... I have 5.6 Sol as advisor/watchdog, and it blocks every turn, I never saw this. Quite a shame, 4.6 wasn't so bad.

4m agoHN ↗

Can you explain your opinion? I'm curious but such vague comments won't convince me.

35m agoHN ↗

Did 4.6 not have an x-high reasoning level? Why are they comparing 4.7 x-high with 4.6 high?

10m agoHN ↗

Not sure on the API side, in Cursor you can always use 4.6 at xhigh.

35m agoHN ↗

Either way, the fact that xAI or SpaceXAI or whatever the name is, I can commend the team behind it on their rapid ascent and progress by being close and or on the frontier in several respects.

12m agoHN ↗

Your comment is like 6 months to a year late.

There for awhile it seemed like we’d have 3 big competitors but then Grok 4.2 or 4.4 was just diabolical while OAI and Claude continued their significant improvements. Grok was/is so bad that I was convinced musk was gonna shut it down and just fund Anthropic compute once they reached their compute agreement.

33m agoHN ↗

( why is the x-axis on the first chart in descending order ? )

32m agoHN ↗

Grok it's really expensive. I'm getting really amazing results using DeepSeek 4.1 Flash for fraction of the price.

24m agoHN ↗

By no one. For the price of 1M token you can get more and with better results with other models.

19m agoHN ↗

What is the most secure way to use this model as someone who is lazy

31m agoHN ↗

Well I "tried it out" I asked it one question, and it gave no answer and said "Sign up to use more!" I don't think I'll be doing that, no.

I can't think of a single dimension grok is winning on (capability, cost, voice), but want to stay open-minded -- anybody want to vouch for its capabilities in any domain?

25m agoHN ↗

It's probably the most aligned (to a single person) model out there!

21m agoHN ↗

For me it works well for agentic coding tasks and terminal/unix/bash (in cursor and grok build); it's also token efficient and cheaper than gpt 5.6. It's def not as good as Fable for me (I haven't used Astra much, can't comment). So it's not the cheapest, not the most capable, but it has a good mix of it for my backend, go, infra work.

The voice is the same AI slop as the others imho.

(This is about Grok 4.6, I didn't test 4.7 yet).

edit: clarified I mean agentic coding tasks

16m agoHN ↗

If you haven't used it, how do you know if it's winning?

I think it's winning on UI for normies (grok bot) and they made some claims about being pareto SOTA (lowest cost per task completed) a while back with 4.6.

I find it to be a perfectly capable model for implementation (there are many in this class--deepseek flash, spark1.3, luna, etc). I find the usage to be very generous w/ supergrok. I find the model to be just fine for 90% of what I want to do, but I use a smarter model to plan complicated things.

23m agoHN ↗

$2/million inout and $6/million output but I couldn't see any pricing information for cached input tokens?

20m agoHN ↗

If the CursorBench 4.0 score diagram is the headline, I read it as "Grok 4.7 xHigh is almost the same as Fable5.1 on low".

Is there a metric for like... time taken when comparing these two? I see score and cost.

If Fable5.1 can knock it out more quickly on low but Grok4.7 might take twice as long to stumble through a problem (and leave behind a bunch of yucky comments or un-needed extra unit tests), are they really comparable?

Or like... the "quality" of the solution? "It works" versus "it's unmaintainable/very messy/hacky".

16m agoHN ↗

Not even close to astra. Astra is something else. It is expensive, but uses way fewer tokens do my tasks.

xAI missed its chance, Ball is on Anthropic's court.

8m agoHN ↗

You'll have to include it in the future, or your benchmark won't be relevant.

For now, I doubt anyone would notice your protest if you didn't announce it.

9m agoHN ↗

Your friendly reminder that Grok is owned and directly steered by the white supremacist guy with the fascist haircut who does nazi salutes and fucked up decades of international order to settle scores for his apartheid south african family. Any amount of using the model supports this.

5m agoHN ↗

Friendly reminder that no sane person believes any of this

5m agoHN ↗

I have to say I'm a little perplexed by HN's perpetual willingness to use grok like it's a normal product made by a normal company.

There's often frustration that every thread related to a Musk company includes a discussion about Musk, but Musk himself caused that by being the only tech founder to actively campaign for Trump. (Zuck, the runner up, didn't do anything even close to this). Every product at every company he owns is hopelessly tethered to his decision to do that, and deserves to be judged on those terms.

Another popular story today is about US rollback of climate regulation [1] under Trump. Some fraction of every cent you spend on Grok goes to support stuff you hate, and yet people get genuinely annoyed when you point it out. Musk and his companies really _are_ special and should be treated as special.

Every thread about a Musk company or product needs a comment like this one. If we had the right values, every comment would look like this one.

[1] https://text.hrw.org/news/2026/09/17/us-revokes-limits-on-po...