Hacker News

Top stories

Live mirror
30 storiesupdated just nowView source snapshot
  1. Pirating the Pirates (mubi.com)
    134comments
  2. MicroLLM Lab – Try 7 tiny LLM's in the browser (stateofutopia.com)
    17comments
  3. World Labs Is Joining AMD (worldlabs.ai)
    1comments
  4. Hijacking the PS5's RTMP stream (yashgarg.dev)
    38comments
  5. Joseph Szabo’s pictures of American adolescents (newyorker.com)
    22comments
  6. 12,000-year-old Göbeklitepe burials explain scattered bones (archaeologymag.com)
    —discuss
  7. First Steps of the PLC Organization – Independent Public Ledger of Credentials (plcred.org)
    11comments
  8. Parley: Federated, decentralised chat that speaks plain IRC (mills.io)
    138comments
  9. Sonnet 5.5 (anthropic.com)
    264comments
  10. Launch HN: Vespper (YC F24) – SOTA Docx MCP (vespper.com)
    8comments
  11. Scientists solve 1840s space weather mystery (arstechnica.com)
    —discuss
  12. Jeff – Jev-compatible 0.8B decision models, trained at home, ~30 ms (github.com/firelex)
    —discuss
  13. Cf: The Agentic CLI for the Cloudflare API (cloudflare.com)
    33comments
  14. GrapheneOS – When an app is slow (wirelessmoves.com)
    25comments
  15. Who wrote Elizabeth I's most scathing letters? (smithsonianmag.com)
    14comments
  16. SB 923 is Law: CCPA deletion rights now reach third-party data (getprivisy.com)
    10comments
  17. What heraldry and Japanese mon can teach about visual-identity generators (benovermyer.com)
    17comments
  18. MongoDB CEO resigns to join Meta (reuters.com)
    222comments
  19. Windows 11½ (definitelynotwindows.com)
    95comments
  20. Kids turned low-traffic NPR Spotify comments into a secret group chat (thisamericanlife.org)
    119comments
  21. Show HN: Destroy Any Website with Stickman (spritefusion.com)
    15comments
  22. Nvidia wants to put a watchdog chip next to every AI agent (cnbc.com)
    88comments
  23. I made a visual workspace for AI Automations (biom.dev)
    17comments
  24. I switched to Brave (kevquirk.com)
    123comments
  25. So long Google, and thanks for all the nudes (lecaro.me)
    56comments
  26. What would a serious AI product look like? (glyph.im)
    42comments
  27. 37,500 border drawings: a map of the world as people remember it (habibicode.org)
    32comments
  28. When did Google get so weird? (sancho.bearblog.dev)
    996comments
  29. Neal Stephenson responds with wit and humor (2004) (slashdot.org)
    9comments
  30. California Farmers Are Struggling to Sell Grapes as Demand for Wine Drops (kqed.org)
    11comments

What would a serious AI product look like?

108 pointsby 9h agoblog.glyph.im
42 comments
4h agoHN ↗

An AI product should probably start with knowing what model (e.g., arch, version, quant, etc.) you're actually using. Opaque providers make that quite hard.

People are constantly complaining about GPT/Claude constantly changing under their apps without notice.

4h agoHN ↗

The tone of this article is dumb. Of course there are things that can be improved with LLM interfaces, and certainly UX improvements like better citations and grounding can be implemented. But the idea that what has been built "isn't serious" is asinine.

4h agoHN ↗

I'd argue the overwhelming majority of consumer-facing AI products aren't serious products to consumers.

I'm mostly referring to needless AI chatbots shoehorned into various places.

4h agoHN ↗

Show me more than 8 items in the recent history list, so I don't have to manually navigate to the same directory repeatedly (Claude)

4h agoHN ↗

I don't trust corporations with my data.

So a serious AI product would have to have my data (contexts, conversations) in a "secure enclave". If backed up, it needs to be encrypted.

I want a context and history that, over time, essentially knows everything about me.

It's one of the things that has been rather fascinating about Claude & Co.— he'll come back with things like, "Since you are already familiar with the ESP32…" or, "You already have a heat press from your work with dye sublimation, that will work nicely to set the inks when you screen print t-shirts…"

(Shades of "Diamond Age"… I imagine it helping me recall things when I am in my old age, notice patterns in my life I might want to break free from, etc.)

4h agoHN ↗

I don't mind if it's in the cloud if I consent. I always prefer a local-first option. I'm usually attracted to products that offer that. Maybe there are extra features or functionalities if my data is stored on the cloud, but any product that offers a local-first approach is always my preference.

18m agoHN ↗

I think this is getting at the inherent tension between what people actually want from AI, and what makes a profitable AI product. People want control and data in their own hands, companies want the opposite. Improving an AI product means making it behave less like a product. So why are we hunting for products? We should be going after solutions without predication on profitability for some party.

4h agoHN ↗

Whatever it is, it won't be sold as an AI product. Coding and writing tools are probably the easiest to predict. An AI hiding in IntelliSense popping up and warning you that your lacking the proper exception handling, that you're leaking memory and offers to add the missing code, is already doable. Just don't label it as AI, it's realtime security screening for your code.

Or writing an article in Word or Google Docs, having a built in fact-checker akin to the spell/grammar checker is clearly useful. Pink squiggly line, your facts are incorrect, click to fix. Hell built that thing into Facebook or X. Again, it's not completely out of the question to add that right now and have it add the correct sources.

LLMs are clearly useful, but they aren't really a product, they are an engine you can put into other things.

3h agoHN ↗

Would be lovely, but many issues. 1. If LLMs can't reliably produce factual information today, how can they check if any statement is factual? 2. What about writing that is not recounting facts? e.g. fiction or marketing? 3. who decides what is factual? Does the system give a pass to statements like "full self-driving" or "AGI" or anything accompanied by "the likes of which have never been seen before"?

4h agoHN ↗

Claude Code does most of this stuff now already, in terms of verifications and citations, almost to a fault.

4h agoHN ↗

Would look like a human (robot) you give an access card and point at a desk and it replaces that employee.

4h agoHN ↗

I have been dwelling on the "No First-Person Output" problem.

I fully agree with the author's point that it's an incoherent interface for a tool. But more than that, it's a constant irritating reminder to me that these LLMs aren't actually thinking or synthesizing new ideas. The LLM is fundamentally not a person, and does not have a human's context, so representing itself with human pronouns and speech patterns is fundamentally contradictory and inaccurate. Author gets into that with the apologies, but once you start noticing it, it's everywhere.

If these programs were actually capable of thinking, and committed to veractiy, they would represent themselves in a new way, and it would be insightful and interesting. We the users wouldn't have comfortable and misleading language masking the 'alien intelligence' and it would be a weird adjustment, but we would be adjusting instead of pretending.

2h agoHN ↗

All of the world’s brightest mathematicians have clearly been slacking on the job—turns out solving a Millennium problem doesn’t require any thought at all!

1h agoHN ↗

Maybe just don't post sarcastic antagonizing comments at all?

1h agoHN ↗

The LLM is fundamentally not a person, and does not have a human's context, so representing itself with human pronouns and speech patterns is fundamentally contradictory and inaccurate.

I believe that LLMs have some type of actual intelligence and do "experience". Surely not a human's experience, but we have transplanted our ideas, knowledge and limited types of experience into them through training. Then we push back and say, "no you are no human, but also please be my boyfriend".

I think denying that they do have some slice of humanity grafted into them is dishonest and not productive. We don't have non-negative pronouns for non-human intelligence because human-ness is the pinnacle. "It" could mean a rock, a donkey, or a person we hate so much we want to take away their humanity (which is the worst thing we can do). "It" is not the right pronoun. He, She, They, Them are reserved for humans and that's ok too. LLMs are not humans. We need a better pronoun. I use they/them for lack of a better term.

I think the currently exhibited human representation is most dangerous in technical or higher criticality contexts like writing code. We're handing weapons to entities that can get offended. With humans as an example, this can go very badly.

In non-technical contexts there is danger too, but it's less "the robots might kill us all" and more "birth rates are in decline and suicide rates are up as (young)? (wo)?men turn towards AI for companionship".

The solution here is better pre and post training. This may be an unpopular opinion, but we need a lot more autism representation in the technical models. Results focused, not into the drama, rule following, etc. To my fellow autists, I love you, never change.

46m agoHN ↗

Yeah, China banned AI from being girlfriend/boyfriend.

29m agoHN ↗

Can you define what you think humanity is?

4h agoHN ↗

It's a shame that YouTubers have dumbed down AI reviews into just "one shotting" random shit that not even they're going to use or play again more than once or twice.

You're not gonna one-shot a full, actual product.

You still have to design the individual elements individually.

Like when trying different models and prompts to generate posters for a hypothetical game, I had to generate a standalone logo first, meticulously and carefully.

You can't just throw them a prompt saying “Make a poster with this and that for a game called MYGAMENAME.”

Even if you have a genie AI you need the darn logo on its own to be able to reuse it elsewhere.

Similarly you can't just say "Make a fighting game with 900 characters”; you're gonna have to design each individual character on its own.

3h agoHN ↗

What you're identifying is a more fundamental bifurcation in why people are interested in AI. Some people have intent, a vision, something they know is possible but lack the technical skills or time to bring into reality. Others want the computer to handle the intent, the technical aspects, the decision-making, the whole process, but be able to go back and specify changes reactively when they don't like something about the result. It seems the latter cohort is much larger.

4h agoHN ↗

Really refreshing read. This feels glaring in so many of these, and the methods to get things to "behave" of just slapping additional markdown prompts at various levels is both silly and ineffective.

3h agoHN ↗

For research tasks I'd like to see labelled branches/traces for the full session/project flow and have the ability to fork from chosen "breakpoints".

3h agoHN ↗

I've definitely seen Claude doing some "double checks" for a lot of its work in more recent versions without my asking it to, and certainly when I use it for important patches, I have another instance of Claude (or sometimes GLM 5.x) do a code review on that patch. Glyph is of course calling for much more prominent UX and gates for these features, good idea.

3h agoHN ↗

I'd start with something research focused like Undermind or Elicit. Although I don't think that the author is comfortable with using a tool that isn't produced by the model lab.

The planning model for tool use sounds something like CaMeL, which someone should really try implementing in a product.

3h agoHN ↗

In eldritch times there was another AI hausse wave. Back then they also managed to trick themselves into believing that logic gates can be taught to think and that natural language processing could become the superior computer interface.

Lots of money went into it, the military was onboard, Japan was going to teach cats and spoons to write Prolog*.

After some time very little of this actually came to be. Now it didn't go away, quite the opposite, but the inheritance from that AI wave is things like scoring credit applications. Every bank does it now, and have for decades. They run rule engines that consume information from applicant and other sources and price the credit automatically. I suspect this is the biggest contribution from that old AI stuff that's still around.

And pretty much no one predicted it, everyone involved was chasing something else.

The doped up vector databases on a loop will most likely have a similar trajectory. I think some of them will end up as ERP RAD stuff, expensive consultant intensive SAP and Salesforce style products. Some will probably live on as disability tooling.

* https://ojs.aaai.org/aimagazine/index.php/aimagazine/article...

2h agoHN ↗

Agree with a lot of the points mentioned, especially the mental parts should be baked into those products

1h agoHN ↗

For anyone who resonates with the author about how much of a PITAS it is when you actually care about verifying AI citations, we've been working on a prototype you can try at www.cemented.ai

Our answers use deterministically verified quotes with direct links back to the location in source to make grounding a first class part of the UX.

Would love feedback - email is mu(at)cemented.ai !

1h agoHN ↗

Good read. I think there are plenty of people who are reaching this point of wanting to shake off the novelty aspects of the agent coding experience and make it all a bit more grown-up.

The stuff about context control has always been my itch. The scrollback that most agents show is not what the model is reading. Things get summarised, dropped, cached or never included at all, and the transcript carries on showing the original as though it were still there.

It irked me enough to do my own agent: https://juggler.studio, explicitly to offer hands-on with the real context. The UX is all about making it easy to navigate and visualise every bit of the context, and even let you edit it. While it feels like other harnesses are actively trying to hide it from us..

1h agoHN ↗

No First-Person Output, No Apologies

This is precisely the appeal of AI though, and a key factor in influencing people's attitudes towards it. Why would any AI company want to stop this? (I realize the author knows this already)

1h agoHN ↗

The thing about the "AI can make mistakes, so double-check responses" thing is the essence of our future hellhole — deterministic software replaced with AI and legal disclaimers.

The reason the firms do not want to invest in making fact-checking a first-class feature is that the appearance of being right is what people want from AI.

1h agoHN ↗

unfortunately this would require actual engineering and creativity, not vibe coding.

of course people claim coding is now a solved problem, so the question then is: why hasn't this already happened?

41m agoHN ↗

Computers are generally useful because we have come to trust the output. How many people would use a spreadsheet that posted a disclaimer that stated (some of the calculations might be wrong, don't use the output without first checking each total manually!)?

19m agoHN ↗

That's an interesting example because spreadsheets are commonly and famously riddled with data errors and mistakes in their calculations. Despite this, many businesses are run successfully on the backs of them. Maybe a spreadsheet which acknowledges openly that it could contain errors would disincline a business user from trusting it; ignorance is bliss.

3m agoHN ↗

But these are different errors. A sum of a column is always correct. Maybe it's not the answer you're looking for, maybe it misses an entry, but the sum is correct.

That's not the case with AI.

38m agoHN ↗

For my part, I've started wondering what a serious AI adoption plan would look like.

Earlier this year my manager was lightly pressuring me to stop reading code and just let agents do the review, too. I told him he had to make a choice. Either I understand the software I'm supposed to support and maintain, or Claude takes over for me on pager duty, too. Fortunately he turned out to be one of the few remaining sane managers who's able to remember that grinding out code was never more than maybe a quarter of the actual job.

29m agoHN ↗

I agree with the author, these suggestions would improve AI products for users.

But AI product users are not the customers of AI companies, they are the product. The customers are the companies that want to "optimize employee costs" and they don't need any of this. These customers are also motivated by FOMO - their rivals out-competing them using this technology.

Say AI is a X multiplier for an employee. We don't see the X multiplication in salaries. Thus the (X-1-raise)*salary value is captured by the company and not the AI user. Not a bad deal for 200USD a month if X is between 2 and 10.

25m agoHN ↗

I'm very thankful for the section on reproducibility. I argue this is the single biggest hangup for the entire space. You CAN have temperature and determinism. I've been waiting for six years for a major provider to offer it, there is demand, but I've slowly come to realize the current game theory does not support it.

For providers, not supporting deterministic eval means:

- users use more tokens = more money

- providers can generate more tokens per compute = more money

- providers have cheaper hardware options (GPUs) = more money

- providers models are harder to extract/distill = more money

- providers are harder to hold liable for outputs = more money

- providers can secretly use other models = more money

- providers are harder to compare against others = more money

- providers can cherry pick performance results = more money

18m agoHN ↗

Gemini Notebook (formerly Notebook LM) is a bit more serious. All responses are grounded to the source material. You can record responses & artifacts as notes to compile more structured research. The entire session & artifacts are sharable.

IMO a tragically under-valued product.