Hacker News

New stories

Live mirror
30 storiesupdated just nowView source snapshot
  1. DBClient: A native Mac database client for 12 engines (I'm the dev)(dbclient.net ↗)
    discuss
  2. I found a better React color picker(npmjs.com ↗)
    discuss
  3. The easiest-to-understand article I've read about Jev(inlevel9.com ↗)
    discuss
  4. Calibre v9.15's "Create your own adventure" writing game(calibre-ebook.com ↗)
    discuss
  5. Chinese CXMT to Match Micron's DRAM Manufacturing Capacity This Year(techpowerup.com ↗)
    discuss
  6. Sonnex – A DAW for macOS(apps.apple.com ↗)
    1comments
  7. The Economic Incentives Behind the Push to Pause AI(vincentschmalbach.com ↗)
    discuss
  8. A Conversation with Arthur Whitney: Can code ever be too terse? (2009)(acm.org ↗)
    discuss
  9. Jev-Leftpad(github.com/f ↗)
    discuss
  10. Oxford Economics Global Cities Index 2026(oxfordeconomics.com ↗)
    discuss
  11. Jev and AI SDK Template by Vercel Labs(github.com/vercel-labs ↗)
    discuss
  12. Show HN: Paste a domain and watch four AI engines judge every page(citegraph.app ↗)
    discuss
  13. Jev at the Branches: The State Machine Is the Agent(stacktoheap.com ↗)
    discuss
  14. Show HN: Sealr – E2EE messenger with per-message rules and screenshot blocking(sealr.chat ↗)
    discuss
  15. Letting Astra decide when to compact its own context(github.com/manuelcecchetto ↗)
    1comments
  16. Holocene Calendar(holocenecalendar.com ↗)
    discuss
  17. Show HN: Deterministic UI testing that checks backend logs, not just the UI(get-verirun.duely.in ↗)
    discuss
  18. Google Home MCP Server(home.google.com ↗)
    discuss
  19. OpenAI's Agentic Software Factory(pragmaticengineer.com ↗)
    discuss
  20. Show HN: I found the free SaaS listing directory(goodsaas.xyz ↗)
    discuss
  21. AI hallucination of Chinese nuclear components almost led to US Military attack(arstechnica.com ↗)
    discuss
  22. Show HN: SnipStash, a snippet keyboard for iOS with no accounts or servers(khancreators.netlify.app ↗)
    discuss
  23. Show HN: Using Qwen to track under-covered news(pressaudit.org ↗)
    1comments
  24. Diogo Almeida: Thoughts on a Typesafe Coding Agent(docs.google.com ↗)
    discuss
  25. Bsdkrun: A Firecracker-style MicroVM for BSD/Linux guests, built on libkrun(github.com/tsirysndr ↗)
    discuss
  26. Unix Year 2038 problem and the art of underestimating(buzzsprout.com ↗)
    discuss
  27. Adversarial examples for fast hash functions(thomasahle.com ↗)
    discuss
  28. Hackathons in the vibe coding era: our token economics experiment(quesma.com ↗)
    1comments
  29. Show HN: Jeeva – A modular trading engine for mid-frequency trading using Jev(github.com/saratangajalaoffl ↗)
    discuss
  30. Use your computer the right way: tmux(chappelle.dev ↗)
    discuss

AI chatbots give wrong answers to financial queries 'most of the time'

88 pointsby 4h agoft.com
39 comments
3h agoHN ↗

Article is just a vague summary of https://www.saturnos.com/report/artificial-authority

Anecdotally, current models seem to be decent at general personal finance principles - certainly better than the majority of personal finance education that people get exposed to unless they seek it out and read a variety of books and sources. But I wouldn't trust them with direct decision making with actual money due to the training lag time on current tax policy, etc.

2h agoHN ↗

certainly better than the majority of personal finance education that people get exposed to unless they seek it out and read a variety of books and sources.

Also, models are now good enough that you can give them chapters from "authoritative" books, and they'll integrate that and come up with better answers even if their "vanilla" answers were average. And they'll tailor stuff to your particular situation. It's funny that the "agentic" stuff is only used in coding mostly, while it can and does work in other fields as well.

As always, you kinda need to check it (at least spot check) but all in all I'd agree it's better than the average stuff you used to find with a quick google search.

1h agoHN ↗

So if you know a book that has the information you need, you just need to upload it into the model to get the right answer.

Isn't that a bit circular - if you already know the authoritative source, why ask a model?

1h agoHN ↗

Because it's faster.

I've done this on different topics - I know the answer is in a particular eBook/PDF/document, but for whatever reason it's not trivial to look it up. The model can do it a lot more quickly than I can, and then I can still verify the accuracy.

1h agoHN ↗

Some of the more widely spread and tolerated LLM outputs seem to be AI slop replacing Journalism/Blog slop. Places where people complained about quality already, but tolerated it if important enough.

So to name some of the more common ones Translation, Summarization, and Reiteration of a source material.

Humans put spin on things, how much you trust a source might not reflect the source's factual accuracy. It might just mean you liked reading it better from one source than another.

58m agoHN ↗

It's basically a fancy context aware ctrl-f to the book.

I use this regularly with RPG manuals. I _know_ the stuff, but don't remember every detail by heart. And just ctrl-f:ing through a Mörk/Pirate Borg -style PDF isn't really productive (they're "artistically" laid out). But I can just ask an AI bot that has the pdf indexed like "how does the medical kit work?" and it'll give me a summary along with the relevant rolls within seconds.

1h agoHN ↗

Important to note is that what is being measured here is the ability of the models not of the chat tools themselves, which combine model completions with other tools that the models can call upon. The mainstream labs already know this about models, it's no secret, and in fact training materials from e.g. Anthropic are at pains to point out that users, or analysts designing workflows, have the reponsibility to ensure the correct tools are used and that human verification takes place at appropriate stages depending on the risk/consequences of the task at hand.

Of course a language-completion model with a training cutoff date won't have up-to-date information on tax rules or the ability to carry out correct numerical calculations, but when you combine that with (in Claude terminology) web search and code execution tools invoked by the chat agent, you immediately have much more reliable results.

17m agoHN ↗

I keep finding that the current harnesses, when encountering syntax that was invalid at training time but is now valid due to new language versions or custom extensions, don't correctly figure out why and assume something is wrong with the codebase or toolchain. I would hate to have that happen with my taxes.

3h agoHN ↗

I would prefer to use agent-assisted python scripts that chatbot.

3h agoHN ↗

Agreed, this works really well for me. Double check the math/python, execute many times without a LLM that can change o

3h agoHN ↗

I like how FT makes me accept cookies from their 46 “technology” (advertising) partners before showing me that the article is behind a paywall anyway.

3h agoHN ↗

You actually like that? I find it kind of annoying.

3h agoHN ↗

No they do not like it, it is a figure of speech to underline how much they do not like it.

3h agoHN ↗

that figure of speech is called sarcasm.

very popular on Earth.

3h agoHN ↗

Single shot or with reasoning enabled? My experience is that reasoning dramatically reduces hallucinations and improves output quality. I don't trust models without it.

2h agoHN ↗

Overall, the best-performing model was Claude Opus 5 on “reasoning” mode, which still made mistakes in 39 per cent of answers.

3h agoHN ↗

These models do pretty well in benchmarks and real world so I'm highly suspicious of this article. Further more, in the original report, the examples of bad answers are from Haiku - at least 7 out of 10. Anyone who knows anything about LLMs know that haiku shouldn't be used for anything pretty much.

There's no reproducible set either. I'm not gonna trust this report.

2h agoHN ↗

Most people[1] interacting with chatbots don't have a paid subscription and they do interact with the free-tier LLMs that are Luna and Haiku, so I still think it's relevant.

[1]: not on HN obviously, but IRL, and probably among FT's readership as well.

1h agoHN ↗

A free claude account with no subscription gets you access to sonnet and I believe uses it by default over haiku

2h agoHN ↗

Now, compare this to a recent story that seemed to claim the opposite:

https://news.ycombinator.com/item?id=49139102

I don't have the time to review the underlying research and decide which one is more correct. My personal biases make me want to believe the current one. Your personal biases may be pulling you in the other direction. How do we make the conversation more intelligent than that?

32m agoHN ↗

One way would be by “[reviewing] the underlying research and [deciding] which one is more correct.”

2h agoHN ↗

Well duh! If it's not using tools to look up the state of the market empirically it's not likely to be accurate financially.

2h agoHN ↗

Given most financial advisors tend to vend out suboptimal advice and steer customers in favour of products they receive a kickback for, I'm happy to be accepting of an unbiased LLM that's trained on bogleheads.org.

2h agoHN ↗

If that's what you want, I'll save you some tokens:

#!/bin/sh

while read question; do echo "Put it into VFIAX"; done

2h agoHN ↗

This is missing a lot of steps like:

- Building an emergency fund

- Budgeting and tracking where your money goes

- Planning and saving for large purchases like cars, homes and life goals

- Optimizing use of tax-advantaged accounts like 401Ks, HSAs, and IRAs

- What to do with ESPPs, RSUs, and options

- How taxes work and how to optimize around them

- Estate planning

1h agoHN ↗

Some people do have complex financial situations. It's not as simple as that.

For example in the UK (and maybe US?) you get tax relief for money you put into your pensions, but there's a limit of £60k/year. Unless you earn a lot (which I do, yeay) when that limit is tapered. Except that you can also use up to 3 years of previously unused allowance. But you have to use this year's first.

Also interest is taxed, but you can put up to £20k/year into an ISA which isn't. And if you still want to avoid some tax you have kids ISA's and even pensions!

Then there are also startup investment schemes that save you some tax. Those seem to be not worth it, but you get the idea - it can be complicated. Especially if you are near one of the many tax/benefit thresholds.

The marginal tax rate in the UK bounces all over the place - it's even technically possible for it to be over 100%!

1h agoHN ↗

While I also practice Bogle's approach from The Little Book of Common Sense Investing, even with this baseline there are some subtleties.

-VFIAX is currently $707/share. Fidelity's FXAIX does not have to be purchased in increments of a share price, and this fund's expenses are lower.

-There are versions of the S&P 500 for taxable accounts that minimize capital gains.

-Vanguard has a total-market index, VTSAX, that is mentioned in the book.

-Vanguard also has a non-U.S. total market fund, VTIAX, that avoid the current CAPE problems of the U.S. market.

Claude is very familiar with Bogle's approach, likely because the pirated book was part of the training set.

36m agoHN ↗

Though you don’t directly say it, your comment strongly implies that you can’t buy fractional shares of $VFIAX. (You can, same as $FXAIX, $VTSAX, etc.)

22m agoHN ↗

There's quite a few implicit assumptions in that.

In my case, I am double-taxed (both Japan and US side) on capital gains. Tax treaties reduce, but not eliminate, the extent of double-taxation.

Many US-based brokers do not allow Americans abroad to purchase mutual funds, so VFIAX is not a choice for me.

Maybe I can go with eMAXIS Slim All Country... Oh, but that is a PFIC under IRS rules and I'd be taxed on unrealized capital gains. So I guess no Japan-equivalents of VT for me. That's fine, I guess I'll just buy VT in my US-based brokerage account; but now I'm in a suboptimal spot with respect to monthly contributions, calculating JPY-denominated income tax on dividends, etc.

I even made an implicit assumption when I said "calculating JPY-denominated income tax on dividends". That assumes your tax status is permanent resident. If your tax status is non-permanent resident, then a decent financial advisor will recognize that only the extent of income remitted to Japan gets taxed, so VT distributing at all isn't an issue (until 5 years later). What should you do before the 5 year threshold is hit? etc. etc.

But yes, if you're born in America and plan to stay within the same state for the rest of your life, then a 100% automated setup that simply deposits $1,000/mo into VFIAX is probably fine. (But keep in mind, to most non-Americans, VOO is not really diversified compared to funds like VT).

Otherwise, there is value in consulting someone (or something) that regularly handles taxes and financial planning.

2h agoHN ↗

So I downloaded that report which of course doesn't contain the most relevant information (the questions) but it contains some examples of wrong answers.

I fed the first question to Grok (which they claimed they tested as well) and it answered it correctly in detail.

I repeated it with another one - again correct answer. I then selected the question they said Grok specifically answered incorrectly and it again answered it correctly.

I am sticking with my first intuition: people are terrible at testing tools and probably wanted them to answer incorrectly/not fully (the questions are constructed in a way to make it difficult as well). They also have vested interest in the conclusion (they are financial advisory firm) so there is that to consider.

People reading ft will now think chat boxes are bad at answering financial questions while they are pretty good at it. Zero consequences for spreading fake news for Financial Times there but good for financial advisors I guess.

51m agoHN ↗

LLMs use random numbers, so a single test won't necessarily match someone else's experience.

1h agoHN ↗

And they hallucinate errors in the millions and struggle with financial data that is in a layout that isn’t in the training data. Ie balance sheet etc.

Been trying to add more AI to my workflow but it just doesn’t work (yet) - not in the same way as vibe coding does

The technical references lookups work though. Looking up regulations etc

1h agoHN ↗

Not a surprise, so do financial advisors: garbage in = garbage out.

1h agoHN ↗

From the official report:

Since LLMs can give different answers to the same question, each question was run five times. That means, each LLM was tested 600 times, and in total over 10,000 questions and answers were assessed.

All models were given the same zero-shot format. They were not given worked examples, previous conversations, hints or an opportunity to correct their answers. This is to make it as similar as possible to a response to a question from consumers.

As for the evaluation itself:

Responses were checked against this (using an LLM-as-a-judge), and was only given a pass if every element was met; otherwise it was assessed as a fail. This all-pass approach was intentionally strict, so that the score measures whether an answer is complete enough to meet the expert legal standard, rather than how many individual points it gets right.

It's just AI slop and it should be taken with a mountain of salt.

23m agoHN ↗

A company selling combined human + AI financial advice finds that AI advice alone is unreliable? Color me surprised.