Hacker News

New stories

Live mirror
30 storiesupdated just nowView source snapshot
  1. Show HN: Cactus Needle 3: 8-29MB automation models can match DeepSeek V4 Flash(cactuscompute.com ↗)
    discuss
  2. Where Awareness Is Not the Problem(medium.com/gurvinder372 ↗)
    discuss
  3. Microsoft exec called AI scraping 'the largest theft of labor in human(techcrunch.com ↗)
    1comments
  4. NASA satellite discovers a new crater on the moon bigger than the Colosseum(cbc.ca ↗)
    discuss
  5. Forget doomsday: The AI hacking crisis is here(axios.com ↗)
    1comments
  6. Base's B20 Bet Is Growing – Could the Next Billion-Dollar Project Be Built Here?(fika.bar ↗)
    discuss
  7. The FAA's plan to fix air traffic? $875M worth of AI(techcrunch.com ↗)
    discuss
  8. Red and Blue America Have Found Something to Agree On: Flock Cameras Must Go(wsj.com ↗)
    discuss
  9. LawConnect(lawconnect.com ↗)
    1comments
  10. Email marketing from Grok Bot official integration(twitter.com/migma_ai ↗)
    discuss
  11. OSU's Bag Man - What was his bag, man?(oregonstater.org ↗)
    discuss
  12. Mistral Hacked(frenchbreaches.com ↗)
    discuss
  13. Dear developers, we are open‑sourcing our dictation app, Blurt(assemblyai.com ↗)
    discuss
  14. Can an AI chatbot save lives by answering texts about pregnancy?(npr.org ↗)
    discuss
  15. Publishers block the robots.txt path their own affiliate redirects live on(disclosed.info ↗)
    discuss
  16. Productivity Effects Across Generations of AI Coding Tools(ssrn.com ↗)
    1comments
  17. Talk to JEV(github.com/mkotlikov ↗)
    discuss
  18. Typing Python, Gradually [video](youtube.com ↗)
    discuss
  19. PPPlayer – An open-source music player built with Flutter(ppplayer.com ↗)
    discuss
  20. We Need to Rewild the Internet (2024)(noemamag.com ↗)
    discuss
  21. Egglog and Equality Saturation in a Production Tensor Compiler(egraphs.org ↗)
    discuss
  22. Fix It in Post(oncemarked.com ↗)
    discuss
  23. Tech continues to be political (2025)(miriamsuzanne.com ↗)
    discuss
  24. SHOW HN: I built the fastest PHP webserver in the world
    1comments
  25. Show HN: A repo-aware architecture checkup for Rails apps(railsbaseline.com ↗)
    discuss
  26. EA Safety(venkateshrao.com ↗)
    discuss
  27. Union Alpha is Unbiased's Pareto(openrouter.ai ↗)
    1comments
  28. Show HN: 500 TB internet index in ClickHouse, with congestion pricing(scry.io ↗)
    discuss
  29. Jev was trained on 100% synthetic data(twitter.com/completeskeptic ↗)
    discuss
  30. Reimagining research papers as interactive and reliable AI agents(nature.com ↗)
    discuss

Show HN: 500 TB internet index in ClickHouse, with congestion pricing

3 pointsby 58m agoscry.io
0 comments
Meet Scry, where I have a 500 TB NVMe internet index (I'm doing my best indexing and normalizing all the intelligence explosion alpha) that you can run ~arbitrary readonly SQL and some of Datalog over, and I handle the problem of resource-contention with congestion-based micro-auction pricing. When there's capacity, the service is free for non-commercial use. ---

Hello. It's 2026, we're training simulated fruit fly brains to play Beat Saber, do we still have to be stuck with internet (re)search as fn: natural language -> black box we can't do anything about -> ranked_list/summary?

There is a long history of people trying to do very fancy things that end up being done in relational databases and a little SQL. There is a gravity to them, a bitter lesson, just like scaling of generalized ml training methods. I mean many, many information products can be built off essentially giant real-time OLAP databases and frontier LLMs writing brilliant SQL+Datalog+vector+Jev etc. queries.

Google Search, Tavily, Exa essentially have the problem of mapping your agents' context you are willing to provide, to a tiny subset of their index. You pay a fixed cost to an extremely hard problem that has a distribution of hardness, which means YOU eat the downsides when they are running out of budgeted compute to help you out.

Their algorithms are opaque to the caller, there's really not much user control, and there's not a serious opportunity to communally improve search recipes, like the lexical+Jev recipes you trust to select bleeding edge AI builders.

Furthermore, search companies aren't even pursuing text-to-SQL anymore (several have talked to me)... they made up their minds during the traumatic 2024 text-to-sql days. They were just too early.

I hope you enjoy. I'm intent on scaling this paradigm on differentiated hardware over much more data, so any compelling use cases or queries I could show off, would be much appreciated!

A quiet thread, for now.Start the conversation on HN ↗