Hacker News

Top stories

Live mirror
30 storiesupdated just nowView source snapshot
  1. Astra for Law(openai.com ↗)
    250comments
  2. Bonsai 2 27B: Near-Lossless Compression in a 9x Smaller Footprint(prismml.com ↗)
    36comments
  3. Bend – A language that blocks AI mistakes via proof, on CPU and GPU(bend-lang.com ↗)
    115comments
  4. Hister: A private search engine for the pages you visit and the files you keep(github.com/asciimoo ↗)
    123comments
  5. Sex, AI, and the Apocalypse(iankduncan.com ↗)
    50comments
  6. Wax motor(wikipedia.org ↗)
    39comments
  7. How to Write with an LLM(sockpuppet.org ↗)
    17comments
  8. Fujitsu launches made-in-Japan next-generation CPU FUJITSU-MONAKA(global.fujitsu ↗)
    180comments
  9. Flet 1.0 – Build cross-platform apps in Python(flet.dev ↗)
    8comments
  10. CrowdSec Source Code Leak(crowdsec.net ↗)
    35comments
  11. Diplodocus, Long Thought Exclusively American, Turns Up in Spain(sci.news ↗)
    7comments
  12. More than 100k people in Japan are now aged 100 or older(bbc.com ↗)
    5comments
  13. The most important product decision is what you don't build(liamnugent.me ↗)
    7comments
  14. Infinite-Parameter LLMs: Generating and Adapting Weights from Live Data(arxiv.org ↗)
    26comments
  15. Rate limits on GitLab.com are changing(about.gitlab.com ↗)
    103comments
  16. How Uber Protects Against Retry Storms(uber.com ↗)
    9comments
  17. How GLM built its own inference infrastructure(z.ai ↗)
    258comments
  18. Why I didn’t sign the Fields medallists’ letter(gowers.wordpress.com ↗)
    248comments
  19. Zettascale (YC S24) Is Hiring ASIC/FPGA Engineers to Build Chips for ASI(zscc.ai ↗)
    discuss
  20. TSMC revealing details about next gen A14 node(mapyourshow.com ↗)
    29comments
  21. The American Religion of Self-Storage Facilities(newyorker.com ↗)
    302comments
  22. How do we prevent mathemathics from devolving into the Medieval Era of secrecy?(mathoverflow.net ↗)
    40comments
  23. One year of sponsored Servo development(servo.org ↗)
    138comments
  24. CCC invites all model citizens to 40C3(ccc.de ↗)
    176comments
  25. Running Ubuntu on the Lenovo IdeaPad Duet(vhaudiquet.fr ↗)
    20comments
  26. Canto: A speech model built for the real world(wisprflow.ai ↗)
    12comments
  27. Computer Reset, Dallas(dfarq.homeip.net ↗)
    discuss
  28. André Weil and the Hodge Conjecture(jiahao116.github.io ↗)
    8comments
  29. Show HN: Snapdrop: Instantly share files between devices. No setup, no signup(snapdrop.me ↗)
    9comments
  30. Launch HN: Skillsync (YC W26) – AI chat sessions made portable across agents
    44comments

Hister – A private, full content search index that you control

497 pointsby 1mo agohister.org
97 comments
26d agoHN ↗

Ohi, author here! Thanks for posting Hister. Feel free to A.M.A.

My first free software search project was Searx, a privacy respecting metasearch engine, but because of the limitations of the metasearch concept, I've decided to take a different approach.

Hister builds a personal search index from pages you visit, bookmarks, browser history, local files, and crawled websites. It stores extracted content with offline result previews, so information remains searchable even when the original page changes or disappears. It supports full text and semantic search, can run entirely on your own machine, and includes a web interface, command line tools, and an MCP endpoint for assistant integrations.

Project page: https://github.com/asciimoo/hister

Tiny read-only demo: https://demo.hister.org/

26d agoHN ↗

interesting you used AGPL 3 licensing, are you planning a hosted version?

would've been great with a more liberal license

26d agoHN ↗

are you planning a hosted version

Not in the near term. Right now I am focused on developing Hister rather than operating a hosted service. There are already plenty of centralized hosted search engines, so my longer term interest is in federation and distributed search. I want to make the core system mature first.

would've been great with a more liberal license

It depends on how do you define liberal. =] I chose AGPLv3+ because I want Hister to remain free software and available to their users.

26d agoHN ↗

it prevents certain innovations to be derived from it but with LLMs I do think these licensing are pretty much moot.

i still think it could've benefited by Apache 2.0 which more or less gets you to your goals

25d agoHN ↗

I agree with AGPLv3, I don't really understand why people have an objection. All it prevents is someone taking your work and making a business out of it.

26d agoHN ↗

it would not! asciimoo, thanks for using AGPL, I've been using hister for months now and it helped me a lot

25d agoHN ↗

Can you give an example of a reasonable thing that you want to do with the Hister source code that the AGPL currently prevents you to do?

26d agoHN ↗

What’s the safari story look like right now?

26d agoHN ↗

Hister builds a personal search index from pages you visit, bookmarks, browser history, local files, and crawled websites.

Immediately interested and will check it out, thank you! I've wanted a "search stuff you've seen online" tool for a long time, but everything seems to be research-oriented or "archive but don't search" or some weird combination that means it's nigh useless to me. I've got decades of bookmarks and archives and I've kinda been stuck grepping them at best (it's rare but I do sometimes want a page I saw once three years ago and I love having that option), while hoping someone would build something better.

One question if ya don't mind, while I explore: any chance of singlefile support? Content-extraction is useful in lots of situations (e.g. wallabag) and it's a great default, but sometimes it fails and sometimes you really do want the page, relatively close to how it actually was. Singlefile does that much better than most, and it does so well enough (and manually-handle-able enough if needed) that I don't feel any desire to switch to WARCs or similar.

Though specifically I'm probably looking for something like "content-extract everything" + "key combo to save singlefile version too" + "upload singlefile archives to backfill / recover". Like 99% of the time content extraction is preferred, and I'm glad to see it... it's just not always enough, and having to go elsewhere for exceptions breaks a lot of the utility.

26d agoHN ↗

Exactly! I had the very same issues before Hister.

One question if ya don't mind, while I explore: any chance of singlefile support?

Yes, partially. Hister can already import HTML files created by SingleFile, but there is no direct integration yet. In the longer term, I would like the SingleFile extension to be able to send snapshots directly to Hister.

26d agoHN ↗

I assume it content-extracts that on upload? I'd really like to move storage into hister too, if possible. That way you could also switch from an extracted view to a "full" view in the UI. Though I assume that'd be fairly simple to build later.

Overall I really like what I'm seeing, it ticks a lot of important boxes for me and it's pleasantly straightforward. Hopefully I'll find time to contribute!

26d agoHN ↗

I assume it content-extracts that on upload?

Hister always stores the original material.

That way you could also switch from an extracted view to a "full" view in the UI.

It isn't even needed, we just need a SingleFile specific extractor (an interface in Hister to parse specific page content and provide custom previews) that provides the full original HTML for the preview panel.

Hopefully I'll find time to contribute!

I'd appreciate it. <3

26d agoHN ↗

Oooh, now I see the extractor-view setting in the UI. Yeah, that's essentially perfect \o/

Thank you again!

26d agoHN ↗

I came here to mention https://github.com/gildas-lormeau/singlefile, I couldn't get it to do what I want so I built my own, for watched domains it pushes a copy of the serialized DOM to a local search database. For structured data, either extract and enrich in the browser or enrich on the server side.

Would Hister support this basic workflow? I'd love to retire my own software.

The next phase was going to move to a recording proxy.

25d agoHN ↗

Singlefile based web pack+download is exactly the next feature I need besides indexing URLs / browser tabs with[0], have it on my to-research for a looong time. I do have a server-side script that can be triggered on doc ingestion as a hook "if data/schema/tab is indexed and path is /ml/* then run script "fetch-url.sh" and file the result into {{doc.path}}" but being able to do that in-browser is more user-friendly.

[0] https://github.com/canvas-ui/canvas/tree/main/apps/browser-e...

26d agoHN ↗

I've been using LinkDing + SingleFile for this. Nice to have another option!

26d agoHN ↗

That is a great combo. My main friction with it is having to manually capture pages I want to keep.

26d agoHN ↗

You should integrate with Karakeep.app, it's only natural that you'd want both a search engine, and a nice "archive" and "article pretty view" features :)

26d agoHN ↗

Very cool. Does it work to connect my phone and laptop to the same firehose?

I often find myself irritated because I read an article on my phone 6 months ago and the history is gone.

26d agoHN ↗

You can host Hister on a home server and access it from multiple devices.

Automatic page capture on mobile currently requires Firefox, since mobile Chrome does not support browser extensions.

25d agoHN ↗

Orion browser hopefully can run the Firefox extension on iOS.

Then joining the mobile device to tailscale network where Hister is hosted on a server on a host in the network … probably this will work at least that is my plan to try.

Thank you for an amazing project. I’m a heavy msgvault user and adding Hister to my browsing will be hugely complementary.

25d agoHN ↗

Can you support kiwi browser mobile? It has extension support and is chromium based.

Happy to contribute the feature if needed

26d agoHN ↗

What would be a typical size of the search index, let's say after 5 years of intense browsing?

26d agoHN ↗

It depends on what you consider intense browsing. An indexed document uses about 100KB on average because Hister stores the full original HTML for offline previews. If storage is a concern, you can disable offline previews, reducing the average document size significantly.

26d agoHN ↗

Do you plan on adding filterable tags to the ui? Rather than just metadata.

26d agoHN ↗

I was thinking about making a "personal data search engine". I have almost 30GB of email archives and a bunch of Google Drive and offline files. I am thinking about downloading it all from the cloud and storing it on a large RAID array, with an indexing and search interface (and possibly a local MCP server).

Would Hister be suitable for this? Can it index mbox files? Would it handle this amount of data? Does it have a search API so I can build an MCP server?

26d agoHN ↗

Would Hister be suitable for this?

The indexer it uses (Bleve) can handle millions of records according to their docs, but sure that Hister would be the best choice for this task. I'd probably use Meilisearch (https://www.meilisearch.com/).

Can it index mbox files?

Not yet.

Does it have a search API so I can build an MCP server?

It has both search API and MCP server endpoints.

25d agoHN ↗

EDIT: I wanted to write "*not* sure" in the previous message.

26d agoHN ↗

This sounds very interesting. Wanted something similar, so going to check it out:)

One thing at first look, like a little feature request ;).

Personally I think it would be very nice if it was possible to index different sites to different search indexs. So you could separate different stuff, like job, different projects you have, and other stuff and then last everything else etc to there on search index db.

So it would be possible to have like "profiles" you could easy choose from. So in the addon, when you click on the hister icon on the toolbar there would be a list of profiles that you could choose where to index the site to. But also a setting to choose which domains and sites that always should index to one search index profils. And all sites that dont is added to a profile is added to the default index instead.

For me this would make it much more easy to find stuff and sort out what Im looking for (for me things get very "thing"/project based what I want to find), so having this..

1 - The "profiles" would become like an important filter. When you search you could easy mark one or more "profiles" which you would search from, and you remove alot of unwanted data automatically, specially when you index every site you visit by automatically.

2 - It would also make it easy if the db gets big over time to remove data that is less important later on and that you dont need to be index anymore.

3 - It would also make it easier to backup only the most important search index DB:s if the DB for the differnt search index:s where in different folders/files. I guess the "default index" that index all sites could easy get big, but there it would be alot of not important data. So would be nice to be able to easy just use a backup program and only backup the only important search databases, to save space.

Thanks

25d agoHN ↗

You can create multiple users for different use-cases. Also, it is possible to automatically assign labels to documents coming from different sources and use these labels to filter results. E.g. using multiple browser profiles where different profiles apply different labels to submitted documents.

26d agoHN ↗

I think its kinda cool, but I dont really get it? So I can gather and search my browser history + most websites i have visited?

26d agoHN ↗

I'd love to be able to use a cloud storage backend (eg. Google Drive) so that my laptop and phone drive space isn't at risk and backups are automatic.

25d agoHN ↗

I'm glad this came up again because the last time is was posted I never got around to installing it on my Unraid. Just did!

25d agoHN ↗

I used to be interested in the prophecies of Nostradamus a lifetime ago. You may be surprised to learn that “Hister” is a name mentioned in those prophecies that is commonly claimed by believers to refer to the future Adolf Hitler.

Academics usually take it to mean the Danube river instead, but in conspiracy/New Age contexts, the association with Hitler is prevalent, and mentioned in several pop culture works, including at least one feature film.

25d agoHN ↗

Literally the first thing I thought of. Is he trying to reference an Antichrist?

25d agoHN ↗

I’m pretty sure it’s just an unintentional name collision. The vast majority of people are unaware of this name’s connotation in Nostradamus, unless they happen to have read or watched some pop culture work that references it.

25d agoHN ↗

Thanks you for making Hister. I have been using it since May and find it immensely useful. It's a piece of mind to know that I always will be able to easily find what I looked at. So I don't feel the need to bookmark things, which I used to do, but then never looked at anyways ;).

25d agoHN ↗

Thanks for your contribution to the Universe! :)

I'm hacking on a in-process DB on top of LMDB+Lance(for now, hilbert space kung-fu with a custom matryoshka embedding setup with separate spatial + temporal + internal and content derived anchors will replace that ~last-century~ last-year tech) + roaring bitmaps as the primary indexing engine with the same or similar goal[0] and will definitely deep-dive into yours.

I'm also trying to index users unstructured documents and workflows(tabs, emails, files, notes, identities etc)

- Organize them into semantically meaningful user or agent created context or directory-like virtual trees (the same photo of a nice kitchen may be surfaced under `/travel/barcelona` and `/arch/interieour/kitches`)

- ..where tree nodes are mapped to bitmaps - `/travel/barcelona` does a fast and cheap `travel` AND `barcelona`, want to "zoom-out" you just go one directory up to `/travel` and see all documents tagged with travel)

- You can use multiple timelines - extract that fancy md-converted en-wiki hf dataset into a wikipedia db dataset + timeline, tag your personal timeline as "personal" - wanna know the zeitgeist of your grandmothers birth date - search for it with timelines personal + wikipedia in layered mode and you'll get everything that happened or was happening during that time.

- You can have long-running stateful query sessions and refine your searches dynamically - search for "winter" and get all documents with a winter scenery or mentioning winter - refine with "nice view" then "laptop" - citing a recent example[1]

- Documents have relations that would be cumbersome to map in a virtual tree structure(worth an experiment due to the zoom-in/out you get with context bitmap trees though) - hence on top of the initial structure you can use graph edges(also powered by bitmaps - as most indexes are)

- All vector queries always run on top of a candidate set you get by the bitmap/bitmap-based filter algebra hence searching through 100k+ docs is usually pretty fast

Anyhow, let me stop here, thank you once again!

[0] https://github.com/canvas-ui/canvas-synapsd (sorry for the sloppy ai readme, no time to resurrect my old one with the updated APIs)

[1] https://demo.cnvs.ai/pub/c/aks6zaf8

25d agoHN ↗

Is the extracted content stored in a way that could be easily converted to markdown? Seems like this could be integrated incredibly well with obsidian as a way to automatically expand your knowledge base, as well as a better search for it.

25d agoHN ↗

The extracted content is stored and displayed as HTML, so converting it to markdown doesn't guarantee lossless transformation. But, the other direction works well: Hister can live track and import markdown files providing full text search and rendered previews for all your files.

23d agoHN ↗

I'm running hister locally, I see in the MCP docs that a third party can query the server, but is there a way to import data as well? I find that I often do research using pi, and would love to save the context to hister to query later on. The closest thing I see is the "hister import file" command, which is great, but inconvenient. If MCP allowed third parties to dump text data in, that would be ideal.

26d agoHN ↗

I tried this out this week and liked it but really wish this project had some form of auth. Opening the contents of every page you’ve ever visited, even to the local network, is not the best idea.

26d agoHN ↗

Hister supports token based, password based, and OIDC/OAuth authentications with optional multi-user handling. Details about user handling can be found here: https://hister.org/docs/user-handling

It also has a "public mode" where anyone can search the indexed content, but only authenticated users can add or modify it.

26d agoHN ↗

It seemed like the public mode was the default when I set it up. If so, that’s a fairly dangerous default as keeping a “clean” history with no secrets leaked seems neigh impossible.

26d agoHN ↗

The default configuration binds only to localhost, and a fresh installation starts with an empty database/index. Could you clarify which specific attack surface you are concerned about in that scenario?

26d agoHN ↗

I’m not concerned about an attack scenario. I’m just saying that using the docker image, if someone (or their agent) isn’t careful, they could expose their browsing history publicly fairly easily. It might just be nice to default to at least a user and pass login rather than just wide open.

26d agoHN ↗

I've set it up today with OIDC AuthN (Keycloak), removed the username/password fields & kept only OIDC, it's working flawless.

Also indexed data is persisted on a per-user basis, so you got this isolation and certainty that your searches will not be polluted by your family's

26d agoHN ↗

Ooh I’ve been thinking about this idea for years, I’m glad someone beat me to it. Guess I have something to play with over the rest of the weekend!

26d agoHN ↗

I set this up a few months ago based on asciimoo's comments on HN, and barely used it at first, but I realized not too long ago that it could be a pretty useful research tool for one of my hobbies (award travel), that revolves around being in the know around various concepts and quirks.

I scraped and imported posts from the blogs I regularly reference for award travel, then hooked it up to OpenCode/Codex as an MCP server and used that corpus for research on those topics. So I can ask things like "has anyone ever mentioned running into this problem before?" [1]

If you have a hobby or working situation that requires you to regularly reference a core set of websites or reference materials, Hister provides almost all the tools out of the box to start a search engine against it. The default datasets they promote include the Python Stlib, MDN and RFC corpus, as an example. [2]

[1]: https://wmchen.com/blog/revisiting-hister/

[2]: https://hister.org/datasets

26d agoHN ↗

Wow, this is a really inspiring use case and blog post. Thanks for sharing it.

What tools or features would Hister need to support your complete search workflow?

26d agoHN ↗

Thank you so much for building this! It's a really awesome piece of work and I'm grateful for your work. I'd love to sponsor you on Github in the near future.

The only thing that I think would be interesting to see is native support for crawling via a sitemap.xml instead of recursively. I worked around this by implementing a basic scraper that fetched pages exclusively from the sitemap.xml to add into Hister.

I think you're already aware of this, but I also experienced some data loss during the import because I was running a concurrent reindex. I clocked it pretty quickly so I didn't think too much of it. [1]

[1] "TODO store new documents in both indexes while running reindex to guarantee not losing any data." @ https://github.com/asciimoo/hister/blob/master/server/indexe...

26d agoHN ↗

Oh this gets my brain spinning! Thanks for the tips

26d agoHN ↗

I'll be honest I wouldn't blame OP for not finding this (I found a species of hister beetle before anything to do with Hitler)

I feel this is a weird spin on the "Clicks to Hitler" game

26d agoHN ↗

I was not using Wikipedia to get there. This is something I immediately recognized from research about World War II and possible prophetic fulfillment. I definitely wasn't playing a game when I came across this information either, but I get your drift.

26d agoHN ↗

Interestingly, when I searched "hister" on English Wikipedia, I was sent to https://en.wikipedia.org/wiki/Danube#Names_and_etymology

Where it says that “Hister” is a variant of its Latin name.

I think the name is lovely with no reason to change. Surely there are few people who would connect this with silly medieval superstition. Perhaps you can take part in wiping out that association.

25d agoHN ↗

I've been on site site for years. I created an account just to comment on this thread. Holy crap, the only time I've ever heard 'Hister' was in a book I read on Nostradamus like 30 years ago.

25d agoHN ↗

Yep. I think I know which book you are thinking of. Maybe a search assistant programmed in Go could be named Go-ring! You know, short for "go" and "webring".

26d agoHN ↗

The name coming from HISTory on STERoids. When I checked the name I found only a beetle and part of the Danube called "Hister".

26d agoHN ↗

I didn't know this link. But, I immediately read Hitler when I saw it. Coming back days later... Nope brain still reads Hitler.

It's a terrible choice of name

25d agoHN ↗

I too read it that way. I think that the problem is in the shape of the word being too similar.

Unlike say Hipster with the descender on the p.

26d agoHN ↗

I LOVE the concept. I will play around with the execution, if it works as described this is a great product.

26d agoHN ↗

Can it import browser bookmarks? It says browser history and a bunch of other bookmark services but not specifically browser bookmarks.

26d agoHN ↗

Is this reply from Gemini? Because I'm used to gemini gaslighting me. The answer is the opposite of my question.

26d agoHN ↗

Doh, sorry, the answer was coming from me who did not read the question properly. It currently cannot import browser bookmarks, but it is a good idea. Added to my TODO.

26d agoHN ↗

Ah sorry didn't mean to be mean, I just wondered if you had automated it haha!

26d agoHN ↗

Time to shameless plug my own easy to host your own content search index

Webtm.io

All open source and small enough to deploy. I deploy to cf webworkers so it’s the only place it’s tested.

One cool thing is we work on iOS, chrome and friends, Firefox and pretty much everywhere. We do require you bring your own LLM though.

26d agoHN ↗

I just finished setting this up yesterday and I'm kind of obsessed with it. I was previously a heavy user of Karakeep, but I hated forgetting to save something and losing it. I also think Hister's semantic search is a better solution than AI generated summaries and tags. Support for local docs is super cool too, I have it set up to index my org notes directory.

26d agoHN ↗

Is the design vibe coded in Codex by any chance? Seeing a lot of similar designs out there and it’s really starting to annoy me.

26d agoHN ↗

The whole design is created by a main contributor and IIRC the style is called neobrutalism.

26d agoHN ↗

Really? The website, you mean? :( I thought to myself, "what a cute, well designed webpage!" when I saw it. Maybe with time this site will stand out like those Claudeslop ones.

25d agoHN ↗

I think that the tells, if they are in fact legitimate, are negligible at best or subtle at worst. It looks like the site uses Tailwind, so if AI was involved it wasn't generating the design from scratch. If it did then I'm impressed because usually it sucks. This is alright.

26d agoHN ↗

Would be really cool to have Zotero library integration / import support!

26d agoHN ↗

I love the idea overall, but something about browser extensions give me a bad taste. Am I overthinking it? I realize the extension is open source, of course.

25d agoHN ↗

If you're on android the extension works for Firefox

25d agoHN ↗

Love to see this. I started something similar a couple years ago but didn't quite have the patience to make the browser extensions really bulletproof.

Planning to contribute significant improvements to the vector search side here.

25d agoHN ↗

This is so cool! For the time being I’m still locked into notion for my handwritten knowledge base, but I love this for incorporating external information.

For the semantic search is there any chunking/processing that happens with the content or do you need to be diligent about having a large embedding context (and/or small content)?

25d agoHN ↗

can it search .pdf content as easy as searching content in txt file?

25d agoHN ↗

- lots of stupid questions to the author from a guy who has no idea about search engines

- let us say I want to index every blog ever listed on HN

- should be a small subset of the 400 billion pages out there on the internet no?

- First I need to gather data, what do you use to load so many webpages rapidly? asyncio with aiohttp in python? are there better options?

- how do you handle proxies? rotation? are there libraries you recommend for this?

- what about pages that use cloudflare? or block your request or present a captcha or a challenge of some kind?

- what are the filetypes you collect? only html or media as well?

- where and in what format do you store all these collected files? flat file storage? duckdb? postgres? hstore? something else?

- what is the frequency at which you refresh each page? once a day? once a week? something else?

- what kind of pre-processing do you use on the collected data? remove extra spaces? special characters? some kind of complex regex pipeline? LLM?

- how do you match the incoming query with processed data? simple text matching? regex? vector embedding match? something else?

22d agoHN ↗

A lot of these can be answered by checking the documentation: https://hister.org.

Others could be answered by examining the source code, as it is open source.

It's good etiquette to check existing documentation and resources before badgering OSS project teams with a lengthy list of questions.

25d agoHN ↗

Thanks, this will pair wonderfully with my Capcat.org project.

25d agoHN ↗

What's the disk space usage like? i.e. average per-day/week additional storage in your usage?

25d agoHN ↗

It’s always been bizarre to me how bad browsers themselves are at searching their own history. This looks great.

25d agoHN ↗

Ister (Latin Hister, Ancient Greek Ἴστρος / Istros) is the ancient classical name used by the Greeks and Romans for the Danube River, specifically referring to its lower course.

25d agoHN ↗

While I love Karakeep, the truth is that I'm using it to solve this exact same problem, even though its project goals are not as well aligned. I would seriously consider giving Hister a try because it seems like a better fit. Bonus points for being written in Go.

I suppose the one thing I need would be an easy way to share links into Hister. Hopefully we will see both Android/iOS mobile apps to send links into Hister and better support for Safari. I wonder what could be possible with the extensions support available on mobile browsers too (Safari, MS Edge, Firefox).