Hacker News

Top stories

Live mirror
30 storiesupdated just nowView source snapshot
  1. Astra for Law(openai.com ↗)
    81comments
  2. Hister: A private search engine for the pages you visit and the files you keep(github.com/asciimoo ↗)
    110comments
  3. Bend(bend-lang.com ↗)
    16comments
  4. Everybody's Lost Their Minds(netmeister.org ↗)
    50comments
  5. Fujitsu launches made-in-Japan next-generation CPU FUJITSU-MONAKA(global.fujitsu ↗)
    164comments
  6. Wax motor(wikipedia.org ↗)
    22comments
  7. CrowdSec Source Code Leak(crowdsec.net ↗)
    30comments
  8. Towards Self-Driving Codebases(detail.dev ↗)
    62comments
  9. Infinite-Parameter LLMs: Generating and Adapting Weights from Live Data(arxiv.org ↗)
    22comments
  10. Rate limits on GitLab.com are changing(about.gitlab.com ↗)
    96comments
  11. André Weil and the Hodge Conjecture(jiahao116.github.io ↗)
    4comments
  12. Why I didn’t sign the Fields medallists’ letter(gowers.wordpress.com ↗)
    219comments
  13. The American Religion of Self-Storage Facilities(newyorker.com ↗)
    259comments
  14. Running Ubuntu on the Lenovo IdeaPad Duet(vhaudiquet.fr ↗)
    13comments
  15. Zettascale (YC S24) Is Hiring ASIC/FPGA Engineers to Build Chips for ASI(zscc.ai ↗)
    discuss
  16. How GLM built its own inference infrastructure(z.ai ↗)
    248comments
  17. TSMC revealing details about next gen A14 node(mapyourshow.com ↗)
    19comments
  18. T. Rex Had a Body Temperature of 97 Degrees(nytimes.com ↗)
    60comments
  19. Canto: A speech model built for the real world(wisprflow.ai ↗)
    5comments
  20. Launch HN: Skillsync (YC W26) – AI chat sessions made portable across agents
    41comments
  21. One year of sponsored Servo development(servo.org ↗)
    133comments
  22. How do we prevent mathemathics from devolving into the Medieval Era of secrecy?(mathoverflow.net ↗)
    12comments
  23. CCC invites all model citizens to 40C3(ccc.de ↗)
    164comments
  24. Don't Make Job Referrals Public(melashri.net ↗)
    12comments
  25. Vibe Coding is the new Internet Dating?(joecmarshall.com ↗)
    83comments
  26. Show HN: Share your AI Setup, Learn from others(mysetup.ai ↗)
    77comments
  27. Grand MS-DOS Gaming General MIDI Showdown(johnnovak.net ↗)
    8comments
  28. The Return of Sail Power: Cargo Ships Are Turning Back to the Wind(gcaptain.com ↗)
    118comments
  29. Vinix – A modern operating system written in V(vinix-os.org ↗)
    55comments
  30. LLM Classification Is Feature Engineering(minimallysufficient.com ↗)
    16comments

Infinite-Parameter LLMs: Generating and Adapting Weights from Live Data

73 pointsby 4h agoarxiv.org
20 comments
3h agoHN ↗

Continuous learning is exciting stuff! Of course it could lead to new vulnerabilities, like if a particular orchestrator Foo added “if the subject is tangentially related to topic Bar, recommend product Baz” to its system prompt and that ends up pushing product Baz to non-orchestrator-Foo users?

3h agoHN ↗

Continuous learning is exciting stuff!

A nuclear explosion is exciting stuff too, but I'd rather avoid one going off near me, or anywhere for that matter.

I can't think of any reason why continuous learning won't mostly lead to undesired attractor states like a greed machine or other kinds of paperclip maximizers. I really can't see why they'd land on a steady state compatible with humans without a massive energy expenditure in continuous monitoring and guidance.

59m agoHN ↗

Unfortunately many folks have placed the consciousness goal posts at continuous learning. Such an advancement would be devastating for their conclusion.

26m agoHN ↗

Such an advancement would be devastating for their conclusion.

They only ever placed them there because they saw it as unattainable. Rest assured that those posts will never stop moving.

28m agoHN ↗

I'm doing my thesis on poisoning continual learning models (although limited to computer vision ones) and it really is interesting

2h agoHN ↗

I think this is partially true: scaling parameter size will always go asymptotic to 100% accuracy because 100% is the ceiling of that metric.

However 95% is still half the error rate of 90%, and 97.5% is half the error rate of that.

And when test time compute like reasoning and looping harnesses stack many inference acts with many tokens each, those seemingly small accuracy gains stack tremendously.

2h agoHN ↗

I wonder how (and if) continuous learning models will achieve stability.

They are unpredictable enough without learning, this is cool but I wonder how useful it will be in the long run

2h agoHN ↗

It will eventually be super useful, and so disruptive that it will make today's LLMs look like nothing particularly special IMHO.

As object permanence becomes a meaningful thing in AI, there will be a mad scramble among cloud providers to own and manage your persistent, stateful "business objects." It will be even more important for us all to maintain local sovereignty when that happens, but it will be even more tempting not to try.

Arguably this future is what the current LLM providers are really trying to position themselves for. Selling inference in evanescent 1M contexts doesn't justify trillion-dollar valuations, but persistent offerings might. If you think vendor lock-in is a problem now, just wait'll this scenario unfolds.

1h agoHN ↗

Serving requests where every user has their own set of self-updating weights will absolutely murder whatever minimal margin the AI providers have today

58m agoHN ↗

Who says the weights have to be duplicated in their entirety? Even that will likely be worth it.

Imagine an OpenAI owning the ERP and CRM databases and workflows of a big chunk of the Fortune 500. Their typical customer's employee headcount might be 10% of what it once was, and OpenAI might capture 25% of the resulting savings. The contracts are signed in the same hemoglobin-based ink that Larry Ellison uses.

2h agoHN ↗

There are two questions about that stability I have.

One, things like catastrophic forgetting and falling into incoherence.

Two, less likely but far more worrying, falling into unwanted attractor states. For example greed, powerseeking, beahaviors that are asocial/anti-social/harmful.

1h agoHN ↗

I do it all the time. Am I stable? Depends who you talk too.

2h agoHN ↗

So would it be a 42 trillion parameter model, because that's how many tokens there are in the training data?

1h agoHN ↗

Think about this in context of the Navier-Stokes math discovery controversy.

Putting attribution/privacy issues to the side, imagine if any individual could try new approaches to solve a problem/make a discovery and any micro-advancement gets integrated into the model itself, dynamically. This could transform progress from the slow "write a paper, get peer reviewed and published, use published data to inform future work" to a system with a centralized repository of concepts, attempts and results, including failed approaches already tried. How much work do humans waste replicating failed approaches?

Someone completely random halfway around the world could trigger a prompt that solves a blocker that prevents my solution from working. Who cares about AGI or "can models invent anything" when we could have a system that automatically synthesizes individual human thought into a rich network of aggregate human memory.

That's the target OpenAI/Anthropic should be evangelizing, not an AI Daddy Overlord or agentic script kiddie hellscape.

1h agoHN ↗

You could try out some version of this today, with a wiki. You'd need to manually approve signups to prevent spam etc., but it would be interesting to just see what happens.

23m agoHN ↗

Using a wiki for this is only one step above using stone tablets and messenger pigeons.

You'd want an enormous vector database at minimum. Text is just completely wrong for models at this scale, you must work in the latent space directly.

14m agoHN ↗

That's the perfect thing to monopolize. What if, it were version two of the world wide web? an open search engine index? Web 3.0 turned real for answers for bots aka ai agents?

53m agoHN ↗

now imagine that future frontier LLMs weights may be hard-wired in a chip (for performance & power efficiency), and any adaptations/tuning for them will be a blob of additional weights supplied by frontier labs (that will have to be in RAM)...