Hacker News

Top stories

Live mirror
30 storiesupdated just nowView source snapshot
  1. Android 17 is the first since 3.x to add new APIs without releasing to the AOSP(grapheneos.social ↗)
    330comments
  2. San Francisco Onion Futures Company(onionfutures.com ↗)
    6comments
  3. Science Is Open Software(jepedersen.dk ↗)
    15comments
  4. SDCC – Small Device C Compiler(sourceforge.net ↗)
    7comments
  5. Cloudflare Quick Tunnels(cloudflare.com ↗)
    266comments
  6. Saving another 100TB of RAM(cloudflare.com ↗)
    58comments
  7. How OpenAI Used Its Own LLMs to Design Its Jalapeño Chip(ieee.org ↗)
    69comments
  8. How to Write with an LLM(sockpuppet.org ↗)
    299comments
  9. Show HN: Seal – Letters and passwords that open for your family after you die(github.com/jasonepage ↗)
    discuss
  10. Why building a Rust LSP is hard(rust-glancer.github.io ↗)
    11comments
  11. Typesafe-computer-use drives a Mac toward a goal for 1/50th of a cent per step(github.com/awlevin ↗)
    2comments
  12. Xcode 27.1 Beta Release Notes(developer.apple.com ↗)
    75comments
  13. Harm Laundering in GPT Models: Gender Discrimination Transformed Rather Than(arxiv.org ↗)
    discuss
  14. The first new cat species discovered in 100 years(nationalgeographic.com ↗)
    80comments
  15. Show HN: Cactus Needle 3: 8-29MB automation models can match DeepSeek V4 Flash(cactuscompute.com ↗)
    81comments
  16. Photon-Emission-Guided Laser Fault Injection Enables RP2350 Secure Debug(ledger.com ↗)
    61comments
  17. OpenJev(openjev.com ↗)
    250comments
  18. The Farnese letter(simonklee.dk ↗)
    5comments
  19. Cache-to-Cache: Direct Semantic Communication Between LLMs (2025)(arxiv.org ↗)
    12comments
  20. Goroutine Leak Profiles(go.dev ↗)
    1comments
  21. LispBM is a concurrent Lisp for microcontrollers with message passing(lispbm.com ↗)
    3comments
  22. Cyclomatic Complexity in C#(ndepend.com ↗)
    16comments
  23. Claude Code now reads AGENTS.md if there is no Claude.md(claude.com ↗)
    208comments
  24. Minimal Phone 2(minimalcompany.com ↗)
    199comments
  25. Show HN: LiveWorld – Every 24/7 YouTube live camera on one globe(liveworld.info ↗)
    32comments
  26. Alibaba open-sources AI model that can detect cancer and nearly 150 conditions(scmp.com ↗)
    9comments
  27. Warez: The Infrastructure and Aesthetics of Piracy (2021)(archive.org ↗)
    39comments
  28. Inside ZCode: Silently uploading your Git history to the cloud(ferstar.org ↗)
    94comments
  29. Flock Offers Employees Buyouts as Customers Flee(wired.com ↗)
    3comments
  30. How SpaceX streamlined the Raptor engine(construction-physics.com ↗)
    55comments

Show HN: Scry, programmable internet search w/ congestion pricing

48 pointsby 1d agoscry.io
22 comments
Meet Scry, a 500 TB NVMe internet index in ClickHouse that you can run ~arbitrary readonly SQL and some of Datalog over, and I handle the problem of resource-contention with congestion-based micro-auction pricing. When there's capacity, the service is free for non-commercial use.

---

Hello. It's 2026, we're training simulated fruit fly brains to play Beat Saber, do we still have to be stuck with internet (re)search as fn: natural language -> black box we can't do anything about -> ranked_list/summary?

There is a long history of people trying to do very fancy things that end up being done in relational databases and a little SQL. There is a gravity to them, a bitter lesson, just like scaling of generalized ml training methods. I mean many, many information products can be built off essentially giant real-time OLAP databases and frontier LLMs writing brilliant SQL+Datalog+vector+Jev etc. queries.

Google Search, Tavily, Exa essentially have the problem of mapping your agents' context you are willing to provide, to a tiny subset of their index. You pay a fixed cost to an extremely hard problem that has a distribution of hardness, which means YOU eat the downsides when they are running out of budgeted compute to help you out.

Their algorithms are opaque to the caller, there's really not much user control, and there's not a serious opportunity to communally improve search recipes, like the lexical+Jev recipes you trust to select bleeding edge AI builders.

Furthermore, search companies aren't even pursuing text-to-SQL anymore (several have talked to me)... they made up their minds during the traumatic 2024 text-to-sql days. They were just too early.

I hope you enjoy. I'm intent on scaling this paradigm on differentiated hardware over much more data, so any compelling use cases or queries I could show off, would be much appreciated!

12h agoHN ↗

This is a very good thing, thanks! Have you talked with any of the smaller search engines like Kagi, Qwant, Brave, Mwmbl, DDG, etc to have this supplement the quality of their results? This seems like a big step towards breaking Google and Bing's dominance in search.

11h agoHN ↗

Please have a 'readable version' option so I don't have to exhaust myself parsing the sites layout. I get that it's unique but most of us just want to work out what you're offering in 5-10 seconds of our time.

11h agoHN ↗

I strongly second this, although I must admit it loaded surprisingly fast for me as I'm on a mobile hotspot in the back of a car.

9h agoHN ↗

HN: This site looks like all the other slop, awful to read.

Also HN: This site is doesn't look like other sites, awful to read.

9h agoHN ↗

I'd bet that no one, not even the site's [human] creators, has ever read that homepage end to end. At best, it might have been handed over to a swarm of reviewer agents.

10h agoHN ↗

Congestion pricing for queries sounds innovative.

10h agoHN ↗

Pricing model is hard to understand at a glance. It uses a term "second of query time" which is not a conventional term and not defined anywhere. Also, all pricing related pages seem to be LLM-generated and are hard to read for a human.

Please, just explain in your own words, how the pricing works, without using made up terms invented by an LLM.

6h agoHN ↗

Hi, yes. New users right now get free credit. When the server isn't burdened, queries are currently ~free (under $.01 per second, when most are well under a second). We're just trying to learn here how to use this thing productively together.

As load increases, the price goes up. There's also exponential egress pricing because Scry is a place where computation happens, not a metaphorical torrent. I'm still tweaking things, balancing between multi-user server load, giving normal active human users lots of priority and no fear in using it hard and deliberately, and fair pricing for agents and automated heavy workflows people set up in the background.

9h agoHN ↗

Cool idea. The pricing model reminds me of my time working in algorithmic trading, haha. I'll try this out for some queries I wanted to run.

I suppose the scraping you're doing is a huge part of your value proposition, but I would like to gently nudge you in the direction of making the datasets available via p2p (e.g. a torrent) like how Wikipedia distributes its snapshots in the spirit of democratizing access to data that is becoming increasingly walled off. Also, I think another potential benefit that kind of bulk sharing would have is relieving the congestion from those doing the equivalent operation to extract data via the querying interface.

4h agoHN ↗

let me know if it was faster and more compositionally expressive than you expected

9h agoHN ↗

Furthermore, search companies aren't even pursuing text-to-SQL anymore (several have talked to me)... they made up their minds during the traumatic 2024 text-to-sql days. They were just too early.

Can you expand on this? Is text-to-sql a deadend? (I personally think it is, having worked on a project at $work. But curious to hear about others experience.)

8h agoHN ↗

Cool but please rewrite the text. Ctrl + F "land" gives 29 matches.

8h agoHN ↗

Scry, keyword action that allows a player to look at a specific number of cards from the top of their library and then arrange those cards in any order, placing any number of them on the bottom of the library and the rest on top.

8h agoHN ↗

How did you scrape reddit comments? Doesn't this require expensive licensing from reddit? How do you handle comments that were deleted by users?

2h agoHN ↗

Fascinating stuff. Curious how you handle the ingestion of data from so many sources cost effectively