Top stories

Live mirror
30 storiesupdated just nowView source snapshot
  1. Introducing System One Models and Jev(typesafe.ai ↗)
    317comments
  2. Apple Reference Image: A New Approach for Verified Photography(security.apple.com ↗)
    56comments
  3. Show HN: An e-ink frame that hears birds and draws them as 1800s illustrations(github.com/arnegiacomo ↗)
    190comments
  4. Negativland, Culture Jamming, and the Art of Making Something New(blog.archive.org ↗)
    16comments
  5. Show HN: I made a flight simulator, except you're just a passenger(inflightsimulator.com ↗)
    24comments
  6. An update on Wayback Machine access(blog.archive.org ↗)
    240comments
  7. Gemini 3.8 Live and 3.8 Live Extended Thinking(blog.google ↗)
    230comments
  8. German Rheinmetall open-sources its Battlesuite connected weapon system protcol(rheinmetall.github.io ↗)
    47comments
  9. Recreating Voodoo Graphics and a Late-1990s Gaming PC on an FPGA(nand2mario.github.io ↗)
    13comments
  10. Building a Linux GPU Driver for the M4 Mac Mini in One Month(codyho.dev ↗)
    133comments
  11. Saving Jet Fuel(marksblogg.com ↗)
    28comments
  12. Jean-Pierre Serre turns 100(st-andrews.ac.uk ↗)
    19comments
  13. We got admin access to Baseten's production GitHub(strix.ai ↗)
    140comments
  14. Stay discoverable in search while disallowing AI training(cloudflare.com ↗)
    27comments
  15. An interactive world map of the stories cultures have told(sunnyguha.com ↗)
    discuss
  16. Chopping up books when they're physically too big(mattkirkland.com ↗)
    144comments
  17. Learning to solve hard problems in RL for LLMs by never giving up(mnoukhov.github.io ↗)
    discuss
  18. Show HN: Capsule – Single-file web apps that save their data into SQLite(withcapsule.app ↗)
    123comments
  19. Sierra digital cameras on the Apple II(colino.net ↗)
    5comments
  20. Doing Everyone Else's Job(yosefk.com ↗)
    3comments
  21. A single firm is behind OpenAI, Anthropic, and Meta hacking scandals(effort.news ↗)
    186comments
  22. Let's make quality the norm again(forbrukerradet.no ↗)
    366comments
  23. WangNet – 1.8 MB, zero-dependency Numberwang adjudication in 11 languages(github.com/graafhenk ↗)
    47comments
  24. Suspected sabotage causes major Netherlands rail disruption(bbc.com ↗)
    405comments
  25. Most people prefer traditional architecture(worksinprogress.news ↗)
    253comments
  26. Show HN: Hacking a $20 4G wireless hotspot into a texting device(bkovac.github.io ↗)
    33comments
  27. The CSS Zen Garden dream, finally shipped(josprague.com ↗)
    75comments
  28. Jiga (YC W21) Is Hiring Product Engineer (Remote/US)(jiga.io ↗)
    discuss
  29. The madman's guide to stamp collecting(wsj.com ↗)
    2comments
  30. Show HN: Pizza Bot – An inbox for AI agents that work in the background(github.com/pizza-bot-app ↗)
    24comments

Stay discoverable in search while disallowing AI training

41 pointsby 2h agoblog.cloudflare.com
27 comments
2h agoHN ↗

Cloudflare, enabling the problem and the solution since, how long has it been?

2h agoHN ↗

(checks watch)

17 years

Fun lava lamp story, though

2h agoHN ↗

A bit too late honestly (?). With so many people who have shifted over to reading AI summaries as a primary search response, those with AI-enabled sites will win by attrition.

There is no going back from this. And the internet is a relatively new phenomenon. Recklessly, blindly applying ads to pages in hopes of generating revenue is a very silly thing to do. Technology with ad blockers and now AI summaries has taken that away. New business models, perhaps actually decent ones are required.

Death of ads everywhere? Good fucking riddance.

Posted from LibreWolf.

2h agoHN ↗

All this attitude does is tear down the only viable income source for independent publishers and demonizes them for trying to make money, while everyone let's huge corporations off the hook for it because "well that's just what they do"

1h agoHN ↗

Advertising in the way it’s done is demonic in and of itself. I don’t care - find a better business model.

48m agoHN ↗

Independent publishers can't afford to run an ad network.

You aren't independent. You work for the BigTech company that serves ads on your site.

2h agoHN ↗

What does this new setting actually do? Does it block their IP ranges too, I hope? As if Meta, for instance, is actually going to respect Accountable, via themselves or their partners, quite frankly is eyebrow raising at best.

2h agoHN ↗

what weirds me out is the analytics

theres no way my index.html page with nothing is getting 10000 hits a day

wtf?

2h agoHN ↗

7 per minute.

I use my high school’s website to test Internet connectivity bc the domain is short and they don’t do a TLS redirect (making it easy to detect WiFi portals).

26m agoHN ↗

I use this too, but my high school url is 8 characters. neverssl.com is 12. :)

22m agoHN ↗

Wish the admin never added the SSL redirect, its literally the namesake lol

I just want basically captive.apple.com with a shorter domain and the webserver not even listening on port 443 at all.

Surprised someone hasn't made this yet, it only requires one spare public IP.

2h agoHN ↗

Accountable mixed-use crawlers remain allowed for search. Every other training crawler is blocked, including the training-only crawlers run by Amazon, Anthropic, Meta, and OpenAI — blocking those does not affect search.

We also categorize the relevant crawlers from Amazon, Anthropic, Meta, and OpenAI as Accountable. These organizations separate their Search and Training crawlers, so Cloudflare can block the Training crawler without affecting search.

I find it difficult to trust that either Meta or OpenAI would use their separate search and training crawlers only for the respective purposes. Their pinky promises have no value, IMO. Both companies are premised on deceptive behaviors.

2h agoHN ↗

I wonder if protocols like Web Bot Auth [1] will see wider adoption. At least as a supported mechanism for those bots which identify themselves. The rest probably still have to be treated with Anubis. In my free time I've recently been experimenting with a Web Bot Auth implementation as an Envoy dynamic module [2] to have a way to define some additional policies for the traffic from bots.

[1] https://datatracker.ietf.org/doc/draft-ietf-webbotauth-https... [2] https://github.com/michalskalski/envoy-web-bot-auth

1h agoHN ↗

Thanks for sharing, it is interesting use case

2h agoHN ↗

I maintain a cloud IP ranges database, and I'm going to test this out.

I have my doubts, though. A formal title like "Accountable" (capitalized) sounds deliberate, but I can't help imagining the renewal email:

"Hey, want to renew your Accountable™ license? Just pinky promise again that you use your IPs for what you say you do."

58m agoHN ↗

"Accountable" is just a fancy word for "pinky promise, but with a label." Nothing stops the data from ending up in a training run once it's already been fetched.

15m agoHN ↗

...and the only way to stop[1] that is by effectively DRM'ing everything, which is a level of dystopia that I don't think even Stallman ever anticipated, nor do I want to happen.

[1] Analog hole and other workarounds aside, naturally.

53m agoHN ↗

"Cloudflare classifies bots by behavior, and a single bot can exhibit more than one behavior."

Is that really true

CF classifies anyone not using a popular browser with Javascript enabled as a "bot"

CF fingerprints www users

As an example, look at CF's Permissions-Policy HTTP header. Then look at IA's. One CDN is advertiser-focused, the other is user-focused

IA = Internet Archive

36m agoHN ↗

As far as I can tell, after months of fighting being DDoSed by Anthropic and OpenAI across 50+ sites - Cloudflare also allows what it considers "good bots" through all of your bot blocking rules, with no option to turn this off unless you pay them money.

18m agoHN ↗

CF classifies anyone not using a popular browser with Javascript enabled as a "bot"

This is my biggest complaint about CF. They are implicitly supporting user-agent discrimination in favour of Big Browser, instead of discriminating on actual behaviour.

...and of course there are already companies running tons of VMs with "officially sanctioned" browser + OS stacks, that can get past all these "protections", for a fee.

"AI bots" is the newest boogeyman they came up with to take away freedom.

46m agoHN ↗

If this admin is serious about AI growth they’d make anti-scrapping illegal.

26m agoHN ↗

I don't think this admin even knows what scraping is. Unless it's "scraping the bottom of the barrel", something they're very familiar with.

9m agoHN ↗

Just make user-agent-discrimination illegal.

9m agoHN ↗

This seems useful for my site. I don't really want competitors' AI systems use our research data to train their models.