New stories

Live mirror
30 storiesupdated just nowView source snapshot
  1. Anthropic's proposed AI watchdog METR has deep ties to Effective Altruism(nypost.com ↗)
    discuss
  2. Hugging Face Disables Offensive Cyber AI Model(huggingface.co ↗)
    discuss
  3. Find every missing music royalty in under 5 minutes (and fix the issues)(usemogul.com ↗)
    discuss
  4. EU Floats Canada Becoming the Bloc's First 'Associate Member'(bloomberg.com ↗)
    discuss
  5. More people are seeking emergency care for gambling and it's mostly men and boys(cbc.ca ↗)
    discuss
  6. Mistral X Mozilla: Private, Multilingual AI Browsing(mistral.ai ↗)
    discuss
  7. Coding Agents Have Converged: Why the SWE-Bench Leaderboard Can No Longer Order(arxiv.org ↗)
    discuss
  8. AI safety beyond the frontier labs: uncensored local models(languageops.com ↗)
    discuss
  9. An updated look for the Raspberry Pi Desktop(raspberrypi.com ↗)
    1comments
  10. Gemini for Go Developers: Building Agents in Go(danicat.dev ↗)
    discuss
  11. iCloud+ Includes Apple TV, Arcade, and Curated Music Stations in 100 Countries(macrumors.com ↗)
    1comments
  12. Omarchy ships its own Linux kernel in v4.0.4(github.com/omacom ↗)
    discuss
  13. Organizing Context in a Multi-Agent Harness(langchain.com ↗)
    discuss
  14. Never-again – so your AI agent stops repeating mistakes you fixed(github.com/malaysherasia-ai ↗)
    discuss
  15. There is NO realistic scenario where AI wipes out all of humanity(twitter.com/richardsocher ↗)
    4comments
  16. Senate Blocks Cryptocurrency Regulation(latimes.com ↗)
    discuss
  17. A backtest can pose a question its own data cannot answer(creativebyzuniga.com ↗)
    discuss
  18. Potemkin Understanding in Large Language Models (2025)(arxiv.org ↗)
    discuss
  19. Text Recognition techniques for premodern Italian and Devanāgarī manuscripts(uniqueatpenn.wordpress.com ↗)
    1comments
  20. A Super-Accurate Clock Using a Tiny Microcontroller(hackaday.com ↗)
    discuss
  21. The structure and lasting impact of the Roman road system(nature.com ↗)
    1comments
  22. Jev means structured output is interesting again(seangoedecke.com ↗)
    discuss
  23. Qwen3.8 Max – Cost per task higher than Astra on Artificial Analysis(artificialanalysis.ai ↗)
    1comments
  24. COBOL dev won .NET hackathon with help from AI(theregister.com ↗)
    1comments
  25. Valeriepieris Circle(wikipedia.org ↗)
    discuss
  26. Show HN: Point at anything in your app, say what you want, get a spec(dhamark.com ↗)
    discuss
  27. Show HN: LaunchPad – I built a Launchpad to replace macOS Apps
    discuss
  28. UK mathematician debunks myth around Parthenon's optical illusions(theguardian.com ↗)
    discuss
  29. The Gap Will Grow(backnotprop.substack.com ↗)
    discuss
  30. The secret formula for Apple's rounded corners (2023)(arun.is ↗)
    discuss

Stay discoverable in search while disallowing AI training

61 pointsby 5h agoblog.cloudflare.com
37 comments
5h agoHN ↗

Cloudflare, enabling the problem and the solution since, how long has it been?

5h agoHN ↗

(checks watch)

17 years

Fun lava lamp story, though

5h agoHN ↗

A bit too late honestly (?). With so many people who have shifted over to reading AI summaries as a primary search response, those with AI-enabled sites will win by attrition.

There is no going back from this. And the internet is a relatively new phenomenon. Recklessly, blindly applying ads to pages in hopes of generating revenue is a very silly thing to do. Technology with ad blockers and now AI summaries has taken that away. New business models, perhaps actually decent ones are required.

Death of ads everywhere? Good fucking riddance.

Posted from LibreWolf.

5h agoHN ↗

All this attitude does is tear down the only viable income source for independent publishers and demonizes them for trying to make money, while everyone let's huge corporations off the hook for it because "well that's just what they do"

4h agoHN ↗

Advertising in the way it’s done is demonic in and of itself. I don’t care - find a better business model.

3h agoHN ↗

Independent publishers can't afford to run an ad network.

You aren't independent. You work for the BigTech company that serves ads on your site.

5h agoHN ↗

What does this new setting actually do? Does it block their IP ranges too, I hope? As if Meta, for instance, is actually going to respect Accountable, via themselves or their partners, quite frankly is eyebrow raising at best.

5h agoHN ↗

what weirds me out is the analytics

theres no way my index.html page with nothing is getting 10000 hits a day

wtf?

5h agoHN ↗

7 per minute.

I use my high school’s website to test Internet connectivity bc the domain is short and they don’t do a TLS redirect (making it easy to detect WiFi portals).

3h agoHN ↗

I use this too, but my high school url is 8 characters. neverssl.com is 12. :)

3h agoHN ↗

Wish the admin never added the SSL redirect, its literally the namesake lol

I just want basically captive.apple.com with a shorter domain and the webserver not even listening on port 443 at all.

Surprised someone hasn't made this yet, it only requires one spare public IP.

2h agoHN ↗

They bullshit big time on that, I have caught them not once.

For a site with proven visitors 100,000 per month CF tell me they saved 90,000 over 1,000,000 visitors (rough numbers).

Sure they know large numbers make people feel good.

5h agoHN ↗

Accountable mixed-use crawlers remain allowed for search. Every other training crawler is blocked, including the training-only crawlers run by Amazon, Anthropic, Meta, and OpenAI — blocking those does not affect search.

We also categorize the relevant crawlers from Amazon, Anthropic, Meta, and OpenAI as Accountable. These organizations separate their Search and Training crawlers, so Cloudflare can block the Training crawler without affecting search.

I find it difficult to trust that either Meta or OpenAI would use their separate search and training crawlers only for the respective purposes. Their pinky promises have no value, IMO. Both companies are premised on deceptive behaviors.

5h agoHN ↗

I wonder if protocols like Web Bot Auth [1] will see wider adoption. At least as a supported mechanism for those bots which identify themselves. The rest probably still have to be treated with Anubis. In my free time I've recently been experimenting with a Web Bot Auth implementation as an Envoy dynamic module [2] to have a way to define some additional policies for the traffic from bots.

[1] https://datatracker.ietf.org/doc/draft-ietf-webbotauth-https... [2] https://github.com/michalskalski/envoy-web-bot-auth

4h agoHN ↗

Thanks for sharing, it is interesting use case

5h agoHN ↗

I maintain a cloud IP ranges database, and I'm going to test this out.

I have my doubts, though. A formal title like "Accountable" (capitalized) sounds deliberate, but I can't help imagining the renewal email:

"Hey, want to renew your Accountable™ license? Just pinky promise again that you use your IPs for what you say you do."

4h agoHN ↗

"Accountable" is just a fancy word for "pinky promise, but with a label." Nothing stops the data from ending up in a training run once it's already been fetched.

3h agoHN ↗

...and the only way to stop[1] that is by effectively DRM'ing everything, which is a level of dystopia that I don't think even Stallman ever anticipated, nor do I want to happen.

[1] Analog hole and other workarounds aside, naturally.

3h agoHN ↗

"Cloudflare classifies bots by behavior, and a single bot can exhibit more than one behavior."

Is that really true

CF classifies anyone not using a popular browser with Javascript enabled as a "bot"

CF fingerprints www users

As an example, look at CF's Permissions-Policy HTTP response header on a site with CF "bot protection", i.e., the "checking your browser" CAPTCHA nonsense (challenges.cloudflare.com). Then look at IA's Permissions-Policy response header. One CDN is advertiser-focused, the other is user-focused

IA = Internet Archive

3h agoHN ↗

As far as I can tell, after months of fighting being DDoSed by Anthropic and OpenAI across 50+ sites - Cloudflare also allows what it considers "good bots" through all of your bot blocking rules, with no option to turn this off unless you pay them money.

7m agoHN ↗

The "good bots" also just happen to be from companies that pay cloudflare a lot of money, I imagine.

3h agoHN ↗

CF classifies anyone not using a popular browser with Javascript enabled as a "bot"

This is my biggest complaint about CF. They are implicitly supporting user-agent discrimination in favour of Big Browser, instead of discriminating on actual behaviour.

...and of course there are already companies running tons of VMs with "officially sanctioned" browser + OS stacks, that can get past all these "protections", for a fee.

"AI bots" is the newest boogeyman they came up with to take away freedom.

3h agoHN ↗

If this admin is serious about AI growth they’d make anti-scrapping illegal.

3h agoHN ↗

I don't think this admin even knows what scraping is. Unless it's "scraping the bottom of the barrel", something they're very familiar with.

3h agoHN ↗

Just make user-agent-discrimination illegal.

3h agoHN ↗

This seems useful for my site. I don't really want competitors' AI systems use our research data to train their models.

2h agoHN ↗

I wouldn’t trust an AI company to honor this as far as I could throw them

2h agoHN ↗

I feel like most of these schemes to categorise data as "public but not really" are ultimately doomed to failure. Even if you could trust every AI company in the world to respect these terms, is there anything stopping someone else indexing the data and selling them the information? I know there's copyright law but they're apparently ignoring that anyway.

Ultimately this reminds me of those really early social media profiles (before people understood privacy settings if they even existed) which would say "If you're not my friend you're not allowed to read this page".

If you don't want your content to end up in some database/archive don't publish it for the whole world to see.

1h agoHN ↗

The irony is that search engines are AI companies now. Telling them 'index me for search but don't train your models' is asking them to split a brain that’s already fully merged.

1h agoHN ↗

What does this really mean though? You can use an LLM to search.

59m agoHN ↗

there have been a few links here on hn about content protection based on Markov’s chain. It might be interesting for Cloudflare to add damage for the scraper that tries to query the page, and not just blocking them.

57m agoHN ↗

“Pre-label your most valuable data for us and we promise to not make it too obvious that we’re training on it”

18m agoHN ↗

"Apple, Google, and Microsoft honor or have committed (in a specified time frame) to honor this setting."

Specified Time Frame ;)