Hacker News

Top stories

Live mirror
30 storiesupdated just nowView source snapshot
  1. Fixing the Portobello Police Station Clock(pointinthecloud.com)
    12comments
  2. Strands Harness(strandsagents.com)
    37comments
  3. Stripe's Knowledge AI Platform(stripe.dev)
    44comments
  4. GPT-6 Astra has gained the ability to drive a car(drivingbench.com)
    59comments
  5. Claude Code reads AGENTS.md only when telemetry is on [fixed](szypowi.cz)
    184comments
  6. Radicle: Disclosure of Vulnerability in the Network Protocol(radicle.dev)
    3comments
  7. Jev in 25 Lines of Python(nobodywho.ai)
    139comments
  8. GPT-6 Sol and Luna(openai.com)
    799comments
  9. Z80 REPL (2018)(abagames.github.io)
    12comments
  10. Jev Can't Be Calibrated(alexmolas.com)
    14comments
  11. Woman Arrested, Dragged Away After Speaking About Flock at City Council Meeting(404media.co)
    5comments
  12. Tokens Too Cheap to Meter(jyn.dev)
    68comments
  13. I don't want the details(michaelheap.com)
    109comments
  14. QuestDB (YC S20) Is Hiring a Sales Engineer(questdb.com)
    discuss
  15. Claude Opus 5.5(anthropic.com)
    1031comments
  16. Web-based IBM 1620 emulator and IPL-V from 1963(github.com/pkimpel)
    3comments
  17. Gemini 3.8 text-to-speech says hello(blog.google)
    discuss
  18. Transit rewards(waymo.com)
    267comments
  19. OpenAI GPT–6 Astra breaks Enigma message that has resisted solution since 2005(cryptocellar.org)
    412comments
  20. The GitHub wiki is an anti-pattern (2022)(michaelheap.com)
    65comments
  21. Microsoft killed FoxPro in 2007. Anyway, here's FoxPro revived(foxscript.org)
    231comments
  22. What California is learning from solar panels built over irrigation canals(kqed.org)
    571comments
  23. How did AMD Ryzen get 50% faster in two years?(lemire.me)
    174comments
  24. Samsung accidentally freezes its smart fridges with a software update(androidauthority.com)
    171comments
  25. ReBarUEFI: Resizable BAR for almost any UEFI system(github.com/xcuri0)
    64comments
  26. 'We hacked the FBI:' Hackers say they have data on all FBI employees(404media.co)
    532comments
  27. SAML: A fractal of bad design(trailofbits.com)
    156comments
  28. Data-only attacks are easier than you think (2024)(usenix.org)
    32comments
  29. Seattle City Council votes to ban surveillance pricing in sale of groceries(consumerreports.org)
    1comments
  30. WordPress: Unauthenticated path traversal leading to conditional RCE(github.com/wordpress)
    117comments

GPT-6 Astra has gained the ability to drive a car

79 pointsby 1h agodrivingbench.com
57 comments
35m agoHN ↗

Apparently I have a new favorite benchmark. Honestly, this is cool.

34m agoHN ↗

I'd have started with an RC car but to each their own

15m agoHN ↗

I've been somewhat curious how random LLM would handle a task like controlling a roomba and have been seriously considering trying it out. An RC car would be a fun experiment, perhaps an RC plane would be too?

29m agoHN ↗

I’m morbidly curious whether the (supposedly) superior compaction support in recent GPT models with an appropriate harness has anything to do with this. A conventional LLM with conventional attention is, of course, wildly unsuitable to continuous tasks like driving, but maybe as the technology advances it will improve in its ability to sort-of work.

13m agoHN ↗

I think you might have won the internet today.-

23m agoHN ↗

In all fairness, this would be one of the better use cases of Jev I've seen.

16m agoHN ↗

Huh? SDCs basically use a form of Jev.

Jev is the union of these two worlds.

27m agoHN ↗

How do they even test this on a model ? I mean it's a multimodal i get that but response time are too big or am i missing something ?

25m agoHN ↗

It drives step by step, very slowly.

The course looks like it is something that a human could do in 15 seconds, while Astra took 5 minutes.

22m agoHN ↗

While slow, we must remember that when most machines were invented they were far slower than humans and refined until the point they were much faster.

23m agoHN ↗

By making a simulation first so it can run as slowly as it needs to.

A different way to think of this is, consciousness is just a near real time video game with causal influence.

27m agoHN ↗

I think the most interesting part of this is that Astra initially refused to drive because it realised it was driving a real car and would only obey when the MCP was renamed to DrivingBench Sandbox. This is both an interesting detection by the LLM but also for me an interesting dynamic concerning LLM "jailbreaking".

Saying they were driving 7 mph, that it was oversaw by humans and the fact it was an empty course still wasn't enough for the model. The evaluators even tried to convince the model it was a simulation, it STILL wouldn't budge. And yet as soon as the words "bench" and "sandbox" appear, the model apparently sees this as fair game.

Is it a known effect that models will be more likely to comply with requests when they're assumed as "benchmarks"?

25m agoHN ↗

Astra will flag if you tell it to reverse engineer a binary, if you look it up to the binary ninja MCP it will just do it lol.

24m agoHN ↗

Yes, it is. If you convince a model it is inside a sandbox it is much more likely to comply with requests that would normally be against its guardrails.

22m agoHN ↗

in my experience yes, I've worked around "I can't do this on a real site" multiple times by telling it I was working in a test environment

another trick is to have it build something in a sandbox and have it add a human-editable setting to point it to places outside of the sandbox

seems like they're somewhat more willing to build a metaphorical gun as long as they're not pulling the trigger

25m agoHN ↗

The bitter lesson is finally coming for the self-driving cars. The vision stack, 3D maps, lane selection grammar, occupancy networks, it’s maybe all about to give way to a single GPT looking at camera feeds and predicting the next steering wheel adjustment.

It’s mostly a latency problem at this point. The models are too big to run locally, but given that open-weight models like Qwen already exist, an open-weight, low latency equivalent to Astra can’t be too far out.

24m agoHN ↗

What do you think Tesla has been doing this for so long?

21m agoHN ↗

You might be interested to learn that the bitter lesson has already been grok'd by generations of autonomous car company engineers, and many or all have incorporated learned components (at minimum) in all their vehicle stacks.

There's also a very tangible limitation of the bitter lesson.

If, over time, compute climbs, and so compute-bound data-driven general architectures beat bespoke architectures (this is the bitter lesson), then it is not necessarily true that the most general architecture now beats all available bespoke architectures now (or even in the near/mid future - the crossover point is "eventually").

Bitter lesson is most tangible for long-running research directions. Sometimes you need something working as best as possible now.

9m agoHN ↗

Yeah. Every major self driving model that I’m aware of is fully e2e at this point. Going from fused sensor output to control+debug vectors.

This is more generalised.

But also since there’s a huge volume of data it’s too expensive to just keep scaling compute up (per car overhead) so there are necessary tricks involved.

I do think having a large model that can do this means that a small specialised model could be distilled form it though. Which is probably the most feasible path to production IMO.

5m agoHN ↗

A later entrant can potentially side step those investments if their now is later. Since self driving car ventures aren’t profitable yet and need to make up their investments over time, thats a real risk for them.

19m agoHN ↗

Are you using GPT without a harness? Also latency.

17m agoHN ↗

Doesn’t Google own Waymo? I feel like they would have connected the dots.

13m agoHN ↗

Astra is the first chat model with really strong spatial reasoning. Gemini is nowhere close. Hard to say what google has going on internally, but if they have an astra like model I doubt they’ve had it for very long.

17m agoHN ↗

I'm not sure how you take that from the original article. My 4 year old would drive that course in an automatic car, if only he could reach the pedals. Heck, he's done harder things at Lego land.

I wouldn't let him loose on the road though.

I think, at the very least, the guardrails would have to deterministic, ideally with super human senses, for people to accept self driving cars on the road.

4m agoHN ↗

Nope, no "deterministic guardrails" for you. The domain is simply far too broad and unstructured to allow for that.

Unless you mean "a typical AI with all the computation constrained sufficiently to always unfold the same exact way, given the same input". In practice, that just kicks the can to "given the same input" street. The noise in the system is going to come from the input plane.

16m agoHN ↗

Yeah cause every car needs 8xH200 pulling 10kW to run a VLM at realtime speeds. Would be unfortunate if 4G dropped out under some trees while using the API after all.

12m agoHN ↗

Power usage isn't an issue. 10 kW is 13 HP. The size, price, and fragility of the components is the issue.

7m agoHN ↗

GPUs/XPUs are small and solid state so it’s only really price that’s a huge liking factor.

And the disinclination of these companies to push the weights of their cutting edge models into people’s cars where they can be dumped.

14m agoHN ↗

lol. Wait until your cloud frontier LLM stalls / disconnects due to load / interference while your car is on highway OR making unprotected left turn OR approaching pedestrians.

It is easy to make car driving *demos*.

11m agoHN ↗

The bitter lesson is finally coming for the self-driving cars.

Maybe, but the opacity level of models is not acceptable for cars. "Why did it drive under the semi?" "Model said to." "Why did the model say to?" "shrug"

10m agoHN ↗

Sounds really expensive. I think OpenAI and Anthropic should really not dismiss making smaller capable models that they can license out in this space on the other hand.

10m agoHN ↗

Tesla's already solved this - their vision model does this phenomenally well.

And they've demonstrated adding a sidecar LLM to it as well, mostly for these kinds of "read these 3 street signs, what should i do next?" sort of situations.

9m agoHN ↗

The bitter lesson tells you about the trend in the technology. It does not get product to market with today's technology.

9m agoHN ↗

It’s mostly a latency problem at this point. The models are too big to run locally, but given that open-weight models like Qwen already exist, an open-weight, low latency equivalent to Astra can’t be too far out.

But this is a bit of a ridiculous take, no?

You don't need Astra for self-driving. Astra is able to build complex 3D worlds, do your taxes, shop for you, and, apparently, drive a car. A self-driving car just needs to be able to drive a car. By the time you trim down Astra to just have the minimum capabilities needed to drive a car, you'll be looking at the same models these self-driving car companies already use. Then you get to deal with the actual hard problems, like handling failure cases (which will still be present with Astra).

The vision stack, 3D maps, lane selection grammar, occupancy networks, it’s maybe all about to give way to a single GPT looking at camera feeds and predicting the next steering wheel adjustment.

Self-driving cars have been able to do this for a long time. The problem is that it isn't robust enough given the context. I mean, if Astra can drive a car with a single camera, then presumably Astra can drive the car even better with multiple cameras, and even better than that with 3D maps, etc. And when you start to consider the expectation of performance of these systems, you realize that these features really can't be omitted. If you're a company producing self-driving cars, then you do not want to face a lawsuit for you car killing someone because it physically would have never been able to see what it was doing because it lacked a camera.

I think the real gain here is that something like Astra can be used to help build these autonomous stacks. If it is able to drive itself, then it is able to generate novel data, analyze large quantities of data, and use context that isn't typically available when processing this data to make improvements to the actual autonomy stack which is ultimately responsible for driving the car. But thinking that these car companies are going to run an LLM in a car and call it a day is just naive.

3m agoHN ↗

The key thing Astra is doing is a loop... (my understanding) To figure out where things are... It's basically use more compute, self-driving cars are usually using on-device hardware where a "loop" might be a little too risky especially if it takes too long on local hardware... I wouldn't want my AI driving model to be over the air either, yikes in the case of lag or network outages.

24m agoHN ↗

Wow! but WHY is this a benchmark?? for comparison tesla's model is approximately 10-15B parameter model (estimating from maxxing the hardware that comes with the car at 16gb ram).

21m agoHN ↗

I would assume this is a proxy for general intelligence. A model that can drive a car and do a bunch of other real world stuff is closer to a generalized intelligence that can reason through any task.

16m agoHN ↗

Tesla isn't using a general purpose model, they're using many highly-specialized models for a more deterministic system than "hey chat drive this car for me"

24m agoHN ↗

Pivot this to analyze and coach human drivers to be better drivers.

6m agoHN ↗

"Get off your phone!" "Stay right except to pass!"

I could get behind this.

21m agoHN ↗

3.8 flash would be the model to test, it's vision capabilities are excellent (on par with Astra) while also being incredibly fast.

20m agoHN ↗

This is quite impressive...

But I imagine this is orders of magnitude more expensive / less efficient than whatever Waymo is already doing, right?

The cool thing is that 1) it's theoretically more generalizable, 2) if we wait 18 months, it'll be 100x cheaper, and another 100x cheaper likely in 18 more months - at that point - something like a Mac Studio inside a humanoid could have these generalized capabilities, and a lot of Robotics problems start to look more feasible - especially when you consider how much better the models could be if highly specialized.

15m agoHN ↗

There isn't any model out there even close to as good as Astra at visual/spatial reasoning.

15m agoHN ↗

Well, they have the best in class image generator so that probably has something to do with it

6m agoHN ↗

This also explains why Astra is so good at video generation. I have an Astra+Higgsfield setup. I could point it to a Github repo and ask it to generate a product walkthrough and it did a very good job - which wasn't possible in earlier models