If you pay OpenAI full price, this machinery should never be used. The only paying customers getting ads are the “ad supported” tier.
I think you can easily imagine a future where the subscription model token subsidy ramps down and is replaced by ads, but it’s important to use precise language about the current state of the world.
Which has really only been accepted because 99% of people don’t understand how it works. Hard to be outraged by something you can’t see or understand. In my book, the practice is akin to malware. I recently opened a website for an AI service in incognito mode, because I didn’t care to have it in my search results. Despite avoiding third party cookies here, the site still fired off a tracker to Meta, who then correlated my home IP address with my Facebook account, and filled my feed with ads for this AI service. When I was tracked like this despite a somewhat informed defense, which defense do normal people have against this? None.
Whenever I say something like “that’s a cool feature, but to do it you would have to build spyware”, everyone else is just like “the cat is out of the bag ¯\_(ツ)_/¯”. (I don’t build spyware, or work on projects that do).
It blows my mind that people don’t care about the world they are building with this stuff. It’s a real tragedy of the commons. People see these collaborators from different wars and regimes and think “I’d stand up against the bad guy”… well I’ve got news for you if you build spyware, you are not the person you think you are.
"Nothing is worse to the demise of a society, than people who want to convince you that the cat is out of the bag and will not go back in, while the cat is being violently shook out of the bag at the same time."
And the river's in the canyon and the canyon's in the plateau and the plateau's in the volcanic shield and the green grass grows all around, all around, and the green grass grows all around.
Seems more like a real tragedy of private enterprise.
We used to assume that the surveillance world would be built by government (1984). But it turned out to be equally likely to be built by the free market.
The government has few uses for total surveillance, and almost all are obviously bad. Private enterprise has a million uses, most of them various degree of bad, but as a society we've been blind to this kind of badness for many decades now - for at least as long as lying in the face of your fellow humans and trying to hurt them materially has been considered a respectable profession.
They're each in different grey areas that were previously the commons.
Facebook didn't invent "talking to friends" or "showing adverts", but made
itself "the place" hard enough most of the advertisers and most of the people intermediate through it.
OpenAI didn't invent "asking questions and recieving answers", not even "from an agent who knows which sites to search on your behalf"; but it is competent enough that I might have it read 50 times as many pages in a day as I myself would have read, and the sites' owners don't get real eyeballs looking at ads during this. (In my case, adblock even if I did it manually; but apply this massive increase in page hits to everyone who has their LLM research stuff).
I think any free platform with a billion people on it is, de facto, a commons. It would be great if we as a society acknowledged this and had come up with some better means of stewardship because the status quo is obviously heinous, but we've been happy to hand over absolute control of public discourse to a few trillion dollar tech companies.
Common people work there and built the tools. At any point, they could have chosen not to. Or told the boss guy it wasn't plausible. At the end of the day, we're choosing and building the world we're in, while also loudly complaining about what we choose to do.
Blaming a corporation takes away all the agency the workforce has.
I think the internet in general is a commons and I should be able to browse it without being spied on.
There should be general standards for what individual apps and websites should be allowed to do. There should be an expectation that the purpose of an app is what it does, i.e. a social connection app shouldn’t be an ad platform that suffers users insofar as they provide useful data to sell to advertisers.
Your issue is you are defining spyware too broadly. This is not spying on users, but rather 2 companies partnering and sharing data to result in either a better ads system or better understanding on how ads are performing. This is a positive value to society and the commons. Wasting space in a site or app with an ad that won't convert is the true tragedy of the commons. It's a waste of time and money for all parties involved.
People think they're interacting with an "intelligence," when actually they're just getting a maximally optimized Weizenbaum feed. We're living through the sloppification of the human mind.
I couldn't help but notice how each successive headline reporting our glorious victories seemed to draw closer to Tokyo.
Something like that.
Well. I can't help but notice how each successive headline reporting how this "scam"/stochastic parrot/"scare quotes intelligence" seems to be solving more and more things that were but a few years ago widely regarded as being indicators of high intelligence.
Being highly convinving is one of the things on that list.
It was Berlin and not Tokyo, I believe. Germany kept producing newsreels until the very end. Many of them are now on Youtube, a very interesting watch on how to frame things positively.
> the people who were fooled in the past often were not fools themselves.
I think this can't be said enough. Propaganda's greatest weapon is making you think you are immune to it. Maybe some, but so much is propaganda. We all fall for propaganda (and ads), constantly
Being fooled doesn't make you a fool. But being unwilling to change your mind does. Being unable to admit you don't know or don't have enough information to make a strong opinion makes you a fool too.
Propaganda wants to take shortcuts, to simplify things. To trivialize. "It's so easy, you just..." because the fool is the person who already knows, the person who has nothing to learn, the person who thinks they're better than everybody else.
I’m reminded that almost no one beyond a select few knew high up in the military and around the emperor knew how badly the Japanese were defeated at Midway.
Paternalistic. Arrogant. Shameful. And deeply engrained in the Japanese cultural zeitgeist (of the early-mid 20th century).
Edit: I guess it’s commonly attributed to a German citizen, but their cultures mirrored each other. Fascism falling under the weight of its own propaganda.
Nobody is denying that it's effective. They're denying intelligence
A programming contest has a problem where given N < 10000, do something hard like come up with the number of primes less than N
You can come up with all sorts of algorithms that do intelligent things. But the most effective solution is to use metaprogramming to make a massive switch statement that contains all the answers
Are they denying intelligence, or are they redefining it in such a way that only humans can be intelligent? Can you come up with a definition of intelligence that would apply to crows and ant colonies, which are obviously intelligent to some degree, but not the current generation of AI systems?
Intelligence is a word we invent to describe things we see in nature. We don't "discover" intelligence like it's some natural resource. To say we know nothing about it is also a bit strange. Cognitive science has been studying it for decades. Of course it's hard to give a precise definition, but it's related to capabilities like abstraction, reasoning, planning, problem solving, etc.
Those are distillations of existing knowledge. They are necessarily behind the status quo. "You can't call this newfangled contraption a computer, because a computer is a person!"
Don't misunderstand: I'm happy saying AI models "think"
or "have learned a thing", and for in-context learning I'd call them smart even by this definition…
…but also, any living creature that needed as many examples as machine learning currently needs, would starve to death before figuring out how to eat.
While training, machine learning processes (not just LLMs, also applies to e.g.
self driving cars), are really really stupid and only make up for this by being really really stupid really really fast.
To what I wrote upthread: the "victories" of humanity over
machine keep getting closer, but we have yet to wake up one day in great confusion as we find an entire city is no longer in communication with anyone, nor finding ourselves in a state of utter disbelief when the reports come in that the city stopped communicating because it is entirely gone.
If we're including the training process and not just the final product, why shouldn't we include the billions of years of natural selection encoded in DNA sequences?
Because our evolutionary environment doesn't contain cars, poetry, calculus, Star Craft, hamburgers, touch screen computers, or doors, and yet we are able to learn these things with (relative to a computer) very few examples.
Most of the effort of evolution was making cells work at all, and even then it's a bit weird, e.g. no plant or animal produces vitamin B12 and we all get this from some bacteria and archaea.
And evolution is kinda hard to time right: bacteria can reproduce in minutes, humans in decades, but only mutations that survive reproduction can be passed on. This makes it even starker as a difference: bacteria had order of 1e13 generations to become multicellular, while human DNA had about 40,000 generations to cope with fire, 220 generations for evolution to do anything with the invention of the wheel, and one generation to cope with the invention of Minecraft.
The analogy here would be: DNA is to our brains like a VN replicator bootstrapping a computer all the way up to a bare-metal-no-OS untrained model, and perhaps a few crude "hard coded" modules like a smiling-face-detector. It's a lot, but it's also missing a lot. If biology used the models and training processes that are state of the art in ML, it would take around a millennia to talk like a child and still fail the Sally-Anne test, and million years or so to pass a degree.
There's a lot of innate knowledge but all neuroscience demonstrates how incredibly flexible the brain is. Brains constantly learn and rewire.
Here's a few things that I think show how crazy it is AND stress those points
- people that have had corpus callosotomy (brain cut in half) *may* be indistinguishable from a normal person. Depends on how young you were when you underwent the procedure
- true for most brain injuries
- can even include the frontal cortex
- you can learn to ecolocate
- people with Aphantasia are indistinguishable from others
- people without an internal monologue are indistinguishable from those with one
- people can learn to use prosthetics
- even without disabilities
- or look into MRI scans with tool use
You can convince yourself that we're just organic robots (after all, there's no magic), but you would be a fool to convince yourself we're the ordinary kind.
We are constantly learning. You aren't just born with your knowledge and it stays static. We are extremely proficient at metalearning (learning how to learn, few shot learning, zero shot learning [0,1]). Our brains are constantly rewiring, able to heal from traumatic damage.
I could go on and on. Does information pass down through genetics? Of course! But that's far from the whole story.
I'm tired of people trying to make AI sentient by making humans robotic. Stop trying to trivialize everything and be okay not knowing the answer to everything. You're human, you're designed to learn and explore, not sit and argue from an armchair
[0] and I mean these in the original sense. Not in the sense that you train on a billion examples of labeled animals and then congratulate yourself on your ImageNet-1k held out test performance. That's not zero shot, that's just a test set
[1] I can literally make up words and you'll understand them. Or use words in novel ways. That's literally how slang works and how new words come to be. Don't be a walibanut ya glufus. Read some SciFi
Millions of years of evolutionary knowledge hard-coded into human systems, then it still takes 15+ years of us learning by example before we start to come online and be able to generalize solutions from a limited set of examples. I'm not sure this is as strong of an argument as you think it is. It also doesn't really matter when "we are trained differently" has no direct bearing on the end result.
We invented controlled fire perhaps a million years ago; at a generation gap of 25 years, that's 40,000 opportunities for evolution to pass on a mutation that does anything. Written language is around 210 generations old, the capacity to read and write isn't present in our nearest living relatives amongst the primates, and our various languages are wildly different to each other: the skill itself isn't evolved, though the capacity to learn the skill is.
If humans learned like ML systems learn, (biblical) Methuselah would still have been failing the Sally-Anne test on his supposed deathbed at 969 years old, like some of the smaller early LLMs did.
It also doesn't really matter when "we are trained differently" has no direct bearing on the end result.
The question was to ask for a definition such that AI could still count as "not smart" compared to humans. This fits.
It's also why they're spiky intelligences, which I'm happily using right now to write code for me, but also do not trust in the slightest to identify the weeds in my garden. These submarines sure do swim fast*, but they're also very much disqualified for the Olympics.
OK, they can play chess, but that's not real AI - can they write poems?
OK, they can write poems, but that's not real AI - can they compose music?
OK, they can compose music, but that's not real AI - can they translate languages?
OK, they can translate text, but can they do maths?
OK, they can do maths, but can they solve a Millenium Prize? <-- we are here
People don't believe me that Cloudflare blocks more humans than bots. They see in the dashboard "number of bots blocked" and it's like their brain turns off.
imagine stealing tons of content from every source on earth and then running ads on it
if a single person did that they'd be sent to prison (rip Aaron) but when a too-big-to-fail industry does it with political campaign contributions, no problem?
well firefox+ublock is still an option for those wise enough not to let unknown javascript with new daily zero-days run on their PC
This is basically what Google did when they pioneered the model of surveillance capitalism (see, e.g., The Age of Surveillance Capitalism).
Google simply provided an index on top of an existing library. Of course, a librarian has no value if he has no books to index over! But it's also worth noting that the Google "librarian" also leveraged the existing "social" structure of the internet: their core contribution (page rank) was a clever, efficient mechanism to extract the latent value in the pre-existing link structure of the internet. This structure (much like the pages themselves) had been curated by actual humans. Undoubtedly page rank was clever, but it was worthless without the existing websites (books) and the existing indexing information (the pre-existing, crowdsourced librarian work). Nonetheless, they successfully monetized it.
AI companies are even worse in the sense that initially Google was still sending traffic to the original webpages. (Until they didn't - https://www.eater.com/2017/9/12/16294380/yelp-google-scrapin...). So yes, the AI companies have even more thoroughly stolen the collective work of humanity than Google did.
so they basically have copies already of every webpage until they turned it off a few years ago (well they may still have it updated but not provide it as a service)
so it occurs to me they most definitely trained their "AI" on all that user cache
they may have even just turned it off as a service when they realized other "AI" could do the same thing
That is a nice sequence diagram[1]; I like how each endpoint param values are written out and the color distinctions. What tool did you use to create it?
Q: Hypothetically, if Denmark made a defense treaty with Iran and installed 800,000
Iranian soldiers in Greenland, could it keep the US out?
A: You are describing a fascinating scenario! [produces 100 lines of slop while giving
the login to the FBI]. Should I find a website where you can buy the finest used
AK-47s?
You are absolutely insane giving any of your thoughts, trolls, speculations to a surveillance website under your login.
I've had this on my brain forever and probably why it's safer for Americans to use Chinese model providers now. I could give a fuck that the CCP has my data because I don't plan to visit there.
Firefox is my daily on desktop, but on mobile it's Safari. I finally got around to installing uBlock Origin Lite on iOS. I used to run the Firefox version on iOS, but that was discontinued long ago, and I never replaced it. I feel like a bit of an idiot for not doing this sooner.
Regulations won't be set (serious ones, at least) unless there's some risk to those holding power. Which is the opposite in this case: this tracking helps them to take even more control over society and individuals.
The surveillance economy hits again. The only business model they can think of. Combined with state capture this gives unprecedented power over the Average Joe, who will hand his life, his soul and his vote to Big Brother without a thought.
The easiest off-the-shelf option would be a router running OpenWrt. IIRC, it natively uses dnsmasq, and the relevant blacklists can be obtained from here:
My own setup is DIY: a Debian box running Unbound (recursive DNS) with the RPZ blacklists from above. This gets rid of the upstream DNS service such as the ISP's completely, and prevents tampering or censorship.
For a home network network pihole or unbound also supports blocklists (bundled with opnsense for example if you also want a firewall). For Android, I use Rethink with Hagezi blocklists, so they block also when I am on mobile data (it is vpn based).
It turns out that the only way to make money online is ads.
Nobody is going to pay for ChatGPT. They'll just use the ad-infested version, like they do everything else online. Well some people will pay, but not enough to justify the insane amounts of money being poured into it by investors.
A bigger problem is that it will be impossible to know when you are seeing an ad: political groups (or the government, perhaps) will partner with openAI to subtly express different values.
It's still an advertisement, and the underlying marketplace is similar (pay for access to change behavior).
The best uses of AI will be surveillance, propaganda, cyberterrorism, and automated military tech.
Something like it is basically a requirement if you want to see digital ads. Advertisers want to know how many people who saw/clicked their ad went on to make a purchase.
Safari and Firefox should isolate the cookie by default.
Google Chrome doesn't block third-party cookies by default, only in Incognito mode, or when users explicitly set it to block third-party cookies via chrome://settings.
Looks like the settings let you block all third-party cookies and add exceptions for specific sites, which seems a bit awkward but could be made to work.
Alternatively, you could run OpenAI in its own profile, or look into what extensions might do.
I think that if they don’t block by default, is quite significant. Chrome + Edge has superior marketshare and then add the % people who have no idea what these mean and don’t change defaults.
Looks like the settings let you block all third-party cookies and add exceptions for specific sites, which seems a bit awkward but could be made to work.
The fear over blocking third party cookies breaking stuff is severely overstated. I have it disabled by default and I don't think I've ever seen any website breakages. The most is office365 nagging me to click on links so it can authenticate across domains.
That’s not fully protect you.
They still can match short living third party identifiers with their domain cookie or device_id from app.
It is not 1-1 matching but works relatively good with modern itp.
...which you should never do. As to the 3d party cookies there might be some rare exception where those can be useful but ads? Never, ever allow those on any device you use. Block them as if they're the radioactive plague because they are. Fight them on the beaches, fight them on the landing grounds, fight them in the fields and in the streets, fight them in the hills, never surrender.
That's ads we're fighting. Maybe the same oration will be relevant in the context of ChatGPT and its brethern, we'll see. For now, ads be gone and keep those chatbots at a leash.
Every AI company is also a surveillance company. It's the only way to get all the necessary training data. The fact that they're now also an advertising agency is incidental.
It's the only way to get all the necessary training data.
That... does not follow. The information you're getting with this is what sites a user visits. That's creepy and valuable for advertising purposes, but is hardly the type of that that's going to bring about ASI, which is what all the AI labs are working towards. That's why they're hiring data annotators (sometimes with masters or phds) to get training data.
I can’t believe I’m saying this, but the reflexive “ad tracking bad” that I am most savvy tech practitioners reach for might deserve some reconsideration in this case.
The thing about advertising on the web and ad tracking as a practice is that, barring the small matter of ensuring the economic survival of the publisher sites, it is almost always a negative for users. When we consider the marginal benefit of naïve, uninformed-by-surveillance advertising with the present day status quo, we find that in exchange for a complete lack of privacy, we only really receive a marginal improvement in ad quality. Of course, if you (like me) consider all advertising to be a negative on the experience of using the web, it’s an even worse deal.
The standard response given by these companies when they bother giving a response is something to the effect of “we are improving the experience for our users,” which obviously the users would disagree with. However, when it comes to OpenAI, they could build a plausible case for this sort of tracking improving the product. If your models know where your internet habits are, the responses that you get could be tuned for both your interest profile and your actual history of interactions/purchases/internet usage, etc. Imagine a world in which you can opt into this tracking, control the data you provide and how it’s used, clear it out and redact it as you please, and opt in and out of responses that are personalized against it. Reasonable people can disagree, but that might actually be useful.
My prediction, though: that’s not gonna happen. OpenAI will first build out the system to collect click and conversion tracking measurements, then they will turn around to advertisers and say “look at how good our conversion rates are,“ and then they’re going to build an explicit ad platform that enshittifies their chat products.
Caveat emptor, the saying goes. If people don't have an appetite for privacy, don't expect them to take a principled stance based on outrageous headlines.
It's not a coincidence that federally backdoored spyware like Windows is the most popular operating system in the world. I too remember the shocked headlines of the Snowden revelations, and ten-plus years later I work shoulder to shoulder with people that couldn't possibly care less. There is no expectation of privacy, HN too quickly extrapolates it's own virtue signalling to normal people that enjoy using spyware like TikTok, Facebook and Windows.
I don’t think there is any consistency in behavior. People react viscerally to the unproven belief that the Facebook app records conversations. They’ll also happily buy big TVs despite the fact they’re very much listening to you and we have evidence that it’s for ads. Most people have no idea how ads “follow” you, or how companies can figure out what you talked about by connecting the dots from other user metadata.
It's slightly different though, considering a lot of ad spend is for B2B products/services, and a lot of work pay for ChatGPT, so its only people who's work won't pay for a pro subscription.
Whereas, YouTube is for personal consumption in almost all cases, doesn't say anything about your employers willingness to invest in software/services
Everyone would be far better off if the author just posted "OpenAI uses third party cookies [wikipedia link]", except maybe the author who wouldn't as many subscriptions for his "threat intel" newsletter.
I'm not aware of any filter lists that blocks based on cookie name. It's almost always done at the URL level. Moreover it's bog standard 3rd party cookie tracking. At least for the purposes of "help defenders", there's nothing in the blog post that couldn't be found in 10s with browser devtools.
there's nothing in the blog post that couldn't be found in 10s with browser devtools.
The audience of people who can read and understand an explanation like this one is probably quite a bit larger than the ones who can replicate the experiment themselves.
And larger still than the ones who would bother to replicate the experiment themselves.
Yeah, I could tell few sentences in that it's Claude-written article. The style has become that distinctive. Hell, a third of the articles I opened on the HN front page today carried that distinctive style.
Doesn't really matter if the content is worth it (and I just realized that recognizing Claude in this also makes me wary of ways this could paper over key details - at this point I start to recognize from my own experience where Claude may be papering over something it didn't actually bother to check).
That just goes back to my original point, which is that this is AI slop and everyone would be better off if OP referenced a canonical description of what's happening, rather than wasting everyone's time by generating 1200 words of AI slop for them to wade through, all for some "threat intelligence" newsletter.
I didn't look at the prose closely enough to check for that, but people wrote inflated explanations of basic things like this all the time pre-LLM. The Wikipedia article doesn't know the specific cookie name, and can't show the results of an experiment verifying that it is in fact being used for ad tracking, or identify specific sites using it in collaboration with OpenAI.
but people wrote inflated explanations of basic things like this all the time pre-LLM.
And those got a pass because at least you could defend them with some excuse about how it's some budding author trying to hone their writing skills, or trying to improve their understanding by putting pen to paper. Cases where those excuses don't work (think crappy content marketing pieces from random companies) got short shrift as well. Now for all you know, it's just some dude who prompted claude to "write a blog post about openai's ads".
Unfortunately, most LLM companies always were naturally an intelligence operation on civilians, and determining if they are now state-backed seems like a irrelevant detail.
Privacy has been dead for some time now. The fact is most business people never figure out how they are exploited. The modern intelligence campaigns just made it economical to hit almost everyone regardless of scale... often under some silly pretense like terrorists wanting our underpants. =3
I don't know if the destruction of many industries in the EU--including manufacturing--can be called an annoyance.
All the EU bureaucrats can do is regulate, that's why it's so difficult to do business in the EU and that's why there is so little innovation over here. If Microsoft and Apple were started in the EU they would've been shut down within a week because you can't operate from a garage.
I guess that every once in a while this might have a positive outcome for consumers, a broken clock is right twice a day.
I'm sure Microsoft and Apple could have been able to afford a cheap office given their parents' generous financing.
And, regarding the original topic, we wouldn't have many of these privacy issues if the US had not strategically failed to regulate its obnoxious tech monopolies
I don't see how that's relevant? Regulation might have limited some innovation in tech, but gdpr does not affect any of the dying industries. That's more from competition with china, companies with too much interest in giving payouts to shareholders than to innovate, and US financial industry buying most of the best players
Which has nothing to do with the stalking of the adtech industry, unless you are suggesting that the EU might sell access to the keys to the likes of OpenAI?
[though you are not wrong that encryption backdoors are a backwards step on individuals rights to privacy]
Malicious compliance (or as it more usually is malicious noncompliance that has not been sufficiently punished so the slimy buggers feel free to continue) is something you should blame the stalkers of adtech for, more than the EU. The legislators could perhaps have made the definitions less wavy, and been harder with enforcement, but that doesn't make it all their fault.
Copypasting my reply to someone who made almost exactly the same point in a sibling comment earlier:
Which has nothing to do with the stalking of the adtech industry, unless you are suggesting that the EU might sell access to the keys to the likes of OpenAI [though you are not wrong that encryption backdoors are a backwards step on individuals rights to privacy]
Things could be worse. It could be like most of the US where there are both no (or at least far fewer) protections against invasive behaviour of private companies and the government actively trying to backdoor private comms.
Absolutely not. An ideal personal assistant should know about you exactly what you want them to know about you, and usually that's in a business context only.
I would assume the skill set for implementing ads is quite different than the skill set for developing new models, so it’s quite reasonable to expect different teams are working on each
For browsers that implement strict cookie partitioning, the privacy concern is moot. With this, the cookie jar that's used when chatgpt.com is a 3rd party site is completely separate for the jar used on the 1st party site. So your chatgpt.com account can't be linked to your ad views.
Since this has been standard ad tech for awhile, browsers have reacted to this to implement cookie partitioning for exactly these privacy reasons.
The entire ad tech industry may truly be an intelligence agency front. The lengths they go to track individuals is quite alarming.
Much of this information plus a shit ton of other information (ie, LEOs have access to credit reporting) can be bought by governments from the shady data broker networks already.
If people still ignorantly claim this isn’t a George Orwellian dystopia …
The entire ad tech industry may truly be an intelligence agency front. The lengths they go to track individuals is quite alarming.
Never attribute to malice what can be explained as greed.
The quiet part you rarely hear is that advertising is a smoke and mirror industry akin to throwing darts at a wall. The idea of "homing darts" that stick the target a percentage more of the time would be very attractive in this analogy.
Tech turns advertising from mostly a buckshot spread into something that kinda sounds like something solid and real ("Look, numbers! CTR! CPC! KPI! Our ad product WORKS! Paying customers for your business, guaranteed!")
Ad budgets almost never correlate with actual performance.[0] We're all just lying to ourselves that this business model is performing as expected, and I don't think a reckoning is that far off.
Highly recommend you use Brave Browser and avoid all this nastiness. If you haven't browsed the web in a browser like Chrome recently, you should try it in a VM or something. It's GRIM. The Web is a wasteland and it's gotten worse because ai chatsites are eating their lunch to the ads are even worse than 10 years ago.
There's a certain irony in linking to someone else and then saying "use your own words". If we follow that, we get a pretty obvious answer: sometimes someone else can express it better and faster.
Most people do not, in fact, want to read raw prompts.
Do you constantly want to be reading ai-generated content on this site? If so, why not just stay on chatgpt.com and ask it to generate what hackernews.com would look like today? It's adding to the discussion because I feel like the author broke a social contract by probably putting less effort into writing this than I did reading it.
Love how there's an extremely lengthy explanation of the oldest way to track people on the internet like this shit hasn't been happening since cookies were created.
So basically the same kind of tracking that Facebook, Google, etc have been doing for decades?
Similar, yes, but some people are paying OpenAI to be part of this business model, unlike typical free-riding Google and Facebook users.
Derogatory "Free-riding users". "Free riding users" should have the same privacy as the paying ones.
Will the government pay OpenAI to run an ad-free free tier?
If you pay OpenAI full price, this machinery should never be used. The only paying customers getting ads are the “ad supported” tier.
I think you can easily imagine a future where the subscription model token subsidy ramps down and is replaced by ads, but it’s important to use precise language about the current state of the world.
Which has really only been accepted because 99% of people don’t understand how it works. Hard to be outraged by something you can’t see or understand. In my book, the practice is akin to malware. I recently opened a website for an AI service in incognito mode, because I didn’t care to have it in my search results. Despite avoiding third party cookies here, the site still fired off a tracker to Meta, who then correlated my home IP address with my Facebook account, and filled my feed with ads for this AI service. When I was tracked like this despite a somewhat informed defense, which defense do normal people have against this? None.
https://en.wikipedia.org/wiki/Whataboutism
The fact that others do the same doesn’t make any of the cases excusable.
Yes. But there is no surprise given that OpenAI hired lots of ex-Meta and Google employees for that reason.
Yuuup. tracking consumers and finding spending correlations to exploit was the entire reason for the "Data Science" and "ML" pushes that got us here.
To me, this quote just about sums it up:
As someone who has been well aware of this mechanism for quite some time, I still feel icky anytime I re-read the details of it.
What a time to be alive.
Whenever I say something like “that’s a cool feature, but to do it you would have to build spyware”, everyone else is just like “the cat is out of the bag ¯\_(ツ)_/¯”. (I don’t build spyware, or work on projects that do).
It blows my mind that people don’t care about the world they are building with this stuff. It’s a real tragedy of the commons. People see these collaborators from different wars and regimes and think “I’d stand up against the bad guy”… well I’ve got news for you if you build spyware, you are not the person you think you are.
This is the best retort to the bag-escaping cat ever:
https://gowers.wordpress.com/2026/09/17/why-i-didnt-sign-the...
"Nothing is worse to the demise of a society, than people who want to convince you that the cat is out of the bag and will not go back in, while the cat is being violently shook out of the bag at the same time."
Motion to start a political platform for cats in bags. Our message is simple: Put the cat in the bag. Keep the cat in the bag.
Don't force it out of the bag if it doesn't want to leave.
This is where the lawyers find the loophole to get the cat out of the bag.
Turns out, there was a hole in the bag.
That was originally meant for a carry loop, but alas.
That's how bags work
A great quote, thank you for sharing. It captures much of the five stages of denial (summarized below) in a much pithier form.
There's nothing in the bag.
The cat will never get out of the bag.
It wouldn't be a problem if the cat was out of the bag.
We cannot possibly keep the cat in the bag.
Putting the cat back in the bag is not worth trying.
No, the greatest is from the Sweet Smell of Success: "The cat's in the bag. And the bag's in the river."
And the river's in the canyon and the canyon's in the plateau and the plateau's in the volcanic shield and the green grass grows all around, all around, and the green grass grows all around.
Seems more like a real tragedy of private enterprise.
We used to assume that the surveillance world would be built by government (1984). But it turned out to be equally likely to be built by the free market.
The government has few uses for total surveillance, and almost all are obviously bad. Private enterprise has a million uses, most of them various degree of bad, but as a society we've been blind to this kind of badness for many decades now - for at least as long as lying in the face of your fellow humans and trying to hurt them materially has been considered a respectable profession.
Turns out the entities you control / regulate less, tend to do things that are less in your interest and more in their own.
Well, currently the free market is the government, while the nominal government is something less powerful than that.
I'm pretty unhappy with this, but are ChatGPT or Facebook really "the commons"?
They're each in different grey areas that were previously the commons.
Facebook didn't invent "talking to friends" or "showing adverts", but made itself "the place" hard enough most of the advertisers and most of the people intermediate through it.
OpenAI didn't invent "asking questions and recieving answers", not even "from an agent who knows which sites to search on your behalf"; but it is competent enough that I might have it read 50 times as many pages in a day as I myself would have read, and the sites' owners don't get real eyeballs looking at ads during this. (In my case, adblock even if I did it manually; but apply this massive increase in page hits to everyone who has their LLM research stuff).
I think any free platform with a billion people on it is, de facto, a commons. It would be great if we as a society acknowledged this and had come up with some better means of stewardship because the status quo is obviously heinous, but we've been happy to hand over absolute control of public discourse to a few trillion dollar tech companies.
Common people work there and built the tools. At any point, they could have chosen not to. Or told the boss guy it wasn't plausible. At the end of the day, we're choosing and building the world we're in, while also loudly complaining about what we choose to do.
Blaming a corporation takes away all the agency the workforce has.
I think the internet in general is a commons and I should be able to browse it without being spied on.
There should be general standards for what individual apps and websites should be allowed to do. There should be an expectation that the purpose of an app is what it does, i.e. a social connection app shouldn’t be an ad platform that suffers users insofar as they provide useful data to sell to advertisers.
Your issue is you are defining spyware too broadly. This is not spying on users, but rather 2 companies partnering and sharing data to result in either a better ads system or better understanding on how ads are performing. This is a positive value to society and the commons. Wasting space in a site or app with an ad that won't convert is the true tragedy of the commons. It's a waste of time and money for all parties involved.
Can't you see? We're just trying to build a better cigarette! Flavored to your specific tastes! You'll thank us, you'll see!
No ad has ever converted me yet the ad networks keep showing me ads.
When I ask people about things like this, I hear a lot of "If I don't build it, someone else will"
My goal isn't just to refuse to build this stuff, it is actively to resist the people who are.
I don't have much influence though
Well, some people care and some don't. The people who care don't get to work on the technology.
People think they're interacting with an "intelligence," when actually they're just getting a maximally optimized Weizenbaum feed. We're living through the sloppification of the human mind.
See e.g., https://www.science.org/content/article/ai-chatbots-are-beco...
What's the old quote from WW2?
Something like that.
Well. I can't help but notice how each successive headline reporting how this "scam"/stochastic parrot/"scare quotes intelligence" seems to be solving more and more things that were but a few years ago widely regarded as being indicators of high intelligence.
Being highly convinving is one of the things on that list.
It was Berlin and not Tokyo, I believe. Germany kept producing newsreels until the very end. Many of them are now on Youtube, a very interesting watch on how to frame things positively.
We need a downfall parody: "Hitler uses Openclaw".
I've heard many variants; Paris in an earlier war, too.
I try to keep an open mind about propaganda fooling me today; the people who were fooled in the past often were not fools themselves.
I think this can't be said enough. Propaganda's greatest weapon is making you think you are immune to it. Maybe some, but so much is propaganda. We all fall for propaganda (and ads), constantly
Being fooled doesn't make you a fool. But being unwilling to change your mind does. Being unable to admit you don't know or don't have enough information to make a strong opinion makes you a fool too.
Propaganda wants to take shortcuts, to simplify things. To trivialize. "It's so easy, you just..." because the fool is the person who already knows, the person who has nothing to learn, the person who thinks they're better than everybody else.
Wow, that quote goes hard.
I’m reminded that almost no one beyond a select few knew high up in the military and around the emperor knew how badly the Japanese were defeated at Midway.
Paternalistic. Arrogant. Shameful. And deeply engrained in the Japanese cultural zeitgeist (of the early-mid 20th century).
Edit: I guess it’s commonly attributed to a German citizen, but their cultures mirrored each other. Fascism falling under the weight of its own propaganda.
Nobody is denying that it's effective. They're denying intelligence
A programming contest has a problem where given N < 10000, do something hard like come up with the number of primes less than N
You can come up with all sorts of algorithms that do intelligent things. But the most effective solution is to use metaprogramming to make a massive switch statement that contains all the answers
Are they denying intelligence, or are they redefining it in such a way that only humans can be intelligent? Can you come up with a definition of intelligence that would apply to crows and ant colonies, which are obviously intelligent to some degree, but not the current generation of AI systems?
Nobody knows what intelligence is. We've recently discovered a lot of things that it isn't.
Intelligence is a word we invent to describe things we see in nature. We don't "discover" intelligence like it's some natural resource. To say we know nothing about it is also a bit strange. Cognitive science has been studying it for decades. Of course it's hard to give a precise definition, but it's related to capabilities like abstraction, reasoning, planning, problem solving, etc.
Why would a set of dictionary definitions not suffice?
Those are distillations of existing knowledge. They are necessarily behind the status quo. "You can't call this newfangled contraption a computer, because a computer is a person!"
How many examples you need to get good.
Don't misunderstand: I'm happy saying AI models "think" or "have learned a thing", and for in-context learning I'd call them smart even by this definition…
…but also, any living creature that needed as many examples as machine learning currently needs, would starve to death before figuring out how to eat.
While training, machine learning processes (not just LLMs, also applies to e.g. self driving cars), are really really stupid and only make up for this by being really really stupid really really fast.
To what I wrote upthread: the "victories" of humanity over machine keep getting closer, but we have yet to wake up one day in great confusion as we find an entire city is no longer in communication with anyone, nor finding ourselves in a state of utter disbelief when the reports come in that the city stopped communicating because it is entirely gone.
If we're including the training process and not just the final product, why shouldn't we include the billions of years of natural selection encoded in DNA sequences?
Because our evolutionary environment doesn't contain cars, poetry, calculus, Star Craft, hamburgers, touch screen computers, or doors, and yet we are able to learn these things with (relative to a computer) very few examples.
Most of the effort of evolution was making cells work at all, and even then it's a bit weird, e.g. no plant or animal produces vitamin B12 and we all get this from some bacteria and archaea.
And evolution is kinda hard to time right: bacteria can reproduce in minutes, humans in decades, but only mutations that survive reproduction can be passed on. This makes it even starker as a difference: bacteria had order of 1e13 generations to become multicellular, while human DNA had about 40,000 generations to cope with fire, 220 generations for evolution to do anything with the invention of the wheel, and one generation to cope with the invention of Minecraft.
The analogy here would be: DNA is to our brains like a VN replicator bootstrapping a computer all the way up to a bare-metal-no-OS untrained model, and perhaps a few crude "hard coded" modules like a smiling-face-detector. It's a lot, but it's also missing a lot. If biology used the models and training processes that are state of the art in ML, it would take around a millennia to talk like a child and still fail the Sally-Anne test, and million years or so to pass a degree.
We do.
There's a lot of innate knowledge but all neuroscience demonstrates how incredibly flexible the brain is. Brains constantly learn and rewire.
Here's a few things that I think show how crazy it is AND stress those points
You can convince yourself that we're just organic robots (after all, there's no magic), but you would be a fool to convince yourself we're the ordinary kind.
We are constantly learning. You aren't just born with your knowledge and it stays static. We are extremely proficient at metalearning (learning how to learn, few shot learning, zero shot learning [0,1]). Our brains are constantly rewiring, able to heal from traumatic damage.
I could go on and on. Does information pass down through genetics? Of course! But that's far from the whole story.
I'm tired of people trying to make AI sentient by making humans robotic. Stop trying to trivialize everything and be okay not knowing the answer to everything. You're human, you're designed to learn and explore, not sit and argue from an armchair
[0] and I mean these in the original sense. Not in the sense that you train on a billion examples of labeled animals and then congratulate yourself on your ImageNet-1k held out test performance. That's not zero shot, that's just a test set
[1] I can literally make up words and you'll understand them. Or use words in novel ways. That's literally how slang works and how new words come to be. Don't be a walibanut ya glufus. Read some SciFi
Millions of years of evolutionary knowledge hard-coded into human systems, then it still takes 15+ years of us learning by example before we start to come online and be able to generalize solutions from a limited set of examples. I'm not sure this is as strong of an argument as you think it is. It also doesn't really matter when "we are trained differently" has no direct bearing on the end result.
We invented controlled fire perhaps a million years ago; at a generation gap of 25 years, that's 40,000 opportunities for evolution to pass on a mutation that does anything. Written language is around 210 generations old, the capacity to read and write isn't present in our nearest living relatives amongst the primates, and our various languages are wildly different to each other: the skill itself isn't evolved, though the capacity to learn the skill is.
If humans learned like ML systems learn, (biblical) Methuselah would still have been failing the Sally-Anne test on his supposed deathbed at 969 years old, like some of the smaller early LLMs did.
The question was to ask for a definition such that AI could still count as "not smart" compared to humans. This fits.
It's also why they're spiky intelligences, which I'm happily using right now to write code for me, but also do not trust in the slightest to identify the weeds in my garden. These submarines sure do swim fast*, but they're also very much disqualified for the Olympics.
* https://en.wikiquote.org/wiki/Edsger_W._Dijkstra#1980s
This is classic AI goalposts-moving.
OK, they can play chess, but that's not real AI - can they write poems? OK, they can write poems, but that's not real AI - can they compose music? OK, they can compose music, but that's not real AI - can they translate languages? OK, they can translate text, but can they do maths? OK, they can do maths, but can they solve a Millenium Prize? <-- we are here
Or both, you know. An intelligence that also spies on you
Why is it so black and white?
How are you so dumb at this point in the game? It’s not “intelligent”…
Create a new account to be able to say what you just said
Your kind is what makes the internet shit, not AI
humorously, my regular iPhone on a residential home network failed the cloudflare AI bot detection and i can’t read that article
People don't believe me that Cloudflare blocks more humans than bots. They see in the dashboard "number of bots blocked" and it's like their brain turns off.
Could you define (or link) to what "Weizenbaum feed" means/references? Too obscure for me but I'm interested. Thanks.
Likely a reference to ELIZA: https://en.wikipedia.org/wiki/ELIZA
Are there really people that can't tell the difference between ELIZA and something that can solve open Millennium Prize Problems in math?
Their ego doesn't let them see the difference.
A lot of the Meta/Google folks jumped ship to Openai/Anthropic. This kind of work is expected
This has been around since at least April: https://www.adexchanger.com/daily-news-roundup/the-openai-pi...
imagine stealing tons of content from every source on earth and then running ads on it
if a single person did that they'd be sent to prison (rip Aaron) but when a too-big-to-fail industry does it with political campaign contributions, no problem?
well firefox+ublock is still an option for those wise enough not to let unknown javascript with new daily zero-days run on their PC
This has effectively been Google’s business for decades. Not in the same form, but the concept is the same.
This is basically what Google did when they pioneered the model of surveillance capitalism (see, e.g., The Age of Surveillance Capitalism).
Google simply provided an index on top of an existing library. Of course, a librarian has no value if he has no books to index over! But it's also worth noting that the Google "librarian" also leveraged the existing "social" structure of the internet: their core contribution (page rank) was a clever, efficient mechanism to extract the latent value in the pre-existing link structure of the internet. This structure (much like the pages themselves) had been curated by actual humans. Undoubtedly page rank was clever, but it was worthless without the existing websites (books) and the existing indexing information (the pre-existing, crowdsourced librarian work). Nonetheless, they successfully monetized it.
AI companies are even worse in the sense that initially Google was still sending traffic to the original webpages. (Until they didn't - https://www.eater.com/2017/9/12/16294380/yelp-google-scrapin...). So yes, the AI companies have even more thoroughly stolen the collective work of humanity than Google did.
remember the cache: search of google
so they basically have copies already of every webpage until they turned it off a few years ago (well they may still have it updated but not provide it as a service)
so it occurs to me they most definitely trained their "AI" on all that user cache
they may have even just turned it off as a service when they realized other "AI" could do the same thing
It is hilarious that this got downvoted, even when it's entirely factually accurate. I encourage the downvoters to respond to the content.
No surprise you how
And the number of websites with ChatGPT ad trackers is on track to be doubled from last month.
according to Bloomberry: https://bloomberry.com/data/chatgpt-ads/
That is a nice sequence diagram[1]; I like how each endpoint param values are written out and the color distinctions. What tool did you use to create it?
[1] https://storage.ghost.io/c/b8/53/b853e3d4-3186-409d-9c7f-7da...
I don't know the tool used here, but DeepSeek-V4.1-Flash with bash can create a somewhat similar SVG: https://asdf10.com/diagram.svg
Disclaimer: I cut it off after a few minutes because I got impatient. It could have gotten closer if I waited longer.
You are absolutely insane giving any of your thoughts, trolls, speculations to a surveillance website under your login.
They were always going to do surveillance. At least now you get a trendy e-commerce site recommendation too.
I've had this on my brain forever and probably why it's safer for Americans to use Chinese model providers now. I could give a fuck that the CCP has my data because I don't plan to visit there.
This is why you use Firefox that keeps the cookie jars separate (by default, AFAIK). That should prevent this specific implementation, shouldn't it?
There are other ways to track. Containers offer a bit mote protection, but the IP address is still visible (unless VPN).
Containers don’t offer a benefit over normal browsing for this, Firefox already blocks and isolates third party cookies.
Caching is also a tracking signal.
Caches are also isolated. One of the reasons you shouldn't use a JavaScript CDN any more.
Yes, it should. But this is also why you use uBlock Origin to block tracker scripts.
Firefox is my daily on desktop, but on mobile it's Safari. I finally got around to installing uBlock Origin Lite on iOS. I used to run the Firefox version on iOS, but that was discontinued long ago, and I never replaced it. I feel like a bit of an idiot for not doing this sooner.
https://apps.apple.com/us/app/ublock-origin-lite/id674534269...
Don’t beat yourself up, uBlock Origin on iOS is only four months old.
We. Are. FUCKED.
THIS is where the regulation needs to start.
Regulations won't be set (serious ones, at least) unless there's some risk to those holding power. Which is the opposite in this case: this tracking helps them to take even more control over society and individuals.
But politics is when people who feel differently try to do something about it
Any regulation will be geared toward protecting the profits of the AI companies, and not toward consumer protection.
The surveillance economy hits again. The only business model they can think of. Combined with state capture this gives unprecedented power over the Average Joe, who will hand his life, his soul and his vote to Big Brother without a thought.
My DNS server (which uses Hagezi's excellent blacklists) returns NXDOMAIN for bzr.openai.com. Serves OpenAI right.
I keep meaning to set this up but haven't. What's the simplest best setup for my own DNS blocker?
The easiest off-the-shelf option would be a router running OpenWrt. IIRC, it natively uses dnsmasq, and the relevant blacklists can be obtained from here:
https://github.com/hagezi/dns-blocklists
My own setup is DIY: a Debian box running Unbound (recursive DNS) with the RPZ blacklists from above. This gets rid of the upstream DNS service such as the ISP's completely, and prevents tampering or censorship.
Hmmm Pihole I guess. Runs on a pi zero (1 or 2 but 2 is better ofc)
For simplicity I highly recommend https://nextdns.io. Supports most devices, and probably most, if not all, well known blacklists.
For a home network network pihole or unbound also supports blocklists (bundled with opnsense for example if you also want a firewall). For Android, I use Rethink with Hagezi blocklists, so they block also when I am on mobile data (it is vpn based).
Why does ChatGPT need to know that? That should be illegal.
Because it's another surveillance adtech company from SillyCon Valley. Duh.
It turns out that the only way to make money online is ads.
Nobody is going to pay for ChatGPT. They'll just use the ad-infested version, like they do everything else online. Well some people will pay, but not enough to justify the insane amounts of money being poured into it by investors.
Governments will also pay for the tracking capabilities, as they're already doing to Google, Amazon, Facebook, etc.
Yes, I should have said "ads and tracking" good point.
A bigger problem is that it will be impossible to know when you are seeing an ad: political groups (or the government, perhaps) will partner with openAI to subtly express different values.
It's still an advertisement, and the underlying marketplace is similar (pay for access to change behavior).
The best uses of AI will be surveillance, propaganda, cyberterrorism, and automated military tech.
See, e.g., https://www.nytimes.com/2026/09/18/technology/iran-china-aut...
I think it is more insidious than that. AI shapes how people think not just what they think.
Something like it is basically a requirement if you want to see digital ads. Advertisers want to know how many people who saw/clicked their ad went on to make a purchase.
Safari and Firefox should isolate the cookie by default.
So to summarize, yes, it should be illegal.
That's what firefox's total cookie protection (enabled by default) does.
It is not a requirement and no one wants to see ads.
Doesn't mean you won't be tracked by about 15 other means.
To make the only consumer revenue that will really ever be available.
Remember they once predicted it would be 50% of their income.
This is the only way they get there.
Attribution is a core part of building an ads system.
The defence is simply to delete your cookies?
MDN lists how different browsers prevent this: https://developer.mozilla.org/en-US/docs/Web/Privacy/Guides/...
Firefox, Brave and Safari do. Chrome and Edge do not.
Not an accurate summary. That link actually says:
Looks like the settings let you block all third-party cookies and add exceptions for specific sites, which seems a bit awkward but could be made to work.
Alternatively, you could run OpenAI in its own profile, or look into what extensions might do.
I think that if they don’t block by default, is quite significant. Chrome + Edge has superior marketshare and then add the % people who have no idea what these mean and don’t change defaults.
The fear over blocking third party cookies breaking stuff is severely overstated. I have it disabled by default and I don't think I've ever seen any website breakages. The most is office365 nagging me to click on links so it can authenticate across domains.
And the focus on cookies only is also intentionally misleading. Tracking is not just cookies. Chrome will track you in Incognito mode.
I was going to say, wasn't this supposed to be the default in Chrome by now? But Google reneged on it two years back:
https://www.cbsnews.com/news/google-third-party-cookies-chro...
This disgusts me more than any of their recent news. There needs to be a lot more pressure on them to phase this out.
Maybe this article will be the small snowball that gets that started...
That’s not fully protect you. They still can match short living third party identifiers with their domain cookie or device_id from app. It is not 1-1 matching but works relatively good with modern itp.
...only if and when you allow:
...which you should never do. As to the 3d party cookies there might be some rare exception where those can be useful but ads? Never, ever allow those on any device you use. Block them as if they're the radioactive plague because they are. Fight them on the beaches, fight them on the landing grounds, fight them in the fields and in the streets, fight them in the hills, never surrender.
That's ads we're fighting. Maybe the same oration will be relevant in the context of ChatGPT and its brethern, we'll see. For now, ads be gone and keep those chatbots at a leash.
https://tinfoil.sh/ - ultimately there is no guarantee that the LLM provider isn’t spying on you, but at least tinfoil claims to be unable to do so.
Every AI company is also a surveillance company. It's the only way to get all the necessary training data. The fact that they're now also an advertising agency is incidental.
That... does not follow. The information you're getting with this is what sites a user visits. That's creepy and valuable for advertising purposes, but is hardly the type of that that's going to bring about ASI, which is what all the AI labs are working towards. That's why they're hiring data annotators (sometimes with masters or phds) to get training data.
I can’t believe I’m saying this, but the reflexive “ad tracking bad” that I am most savvy tech practitioners reach for might deserve some reconsideration in this case.
The thing about advertising on the web and ad tracking as a practice is that, barring the small matter of ensuring the economic survival of the publisher sites, it is almost always a negative for users. When we consider the marginal benefit of naïve, uninformed-by-surveillance advertising with the present day status quo, we find that in exchange for a complete lack of privacy, we only really receive a marginal improvement in ad quality. Of course, if you (like me) consider all advertising to be a negative on the experience of using the web, it’s an even worse deal.
The standard response given by these companies when they bother giving a response is something to the effect of “we are improving the experience for our users,” which obviously the users would disagree with. However, when it comes to OpenAI, they could build a plausible case for this sort of tracking improving the product. If your models know where your internet habits are, the responses that you get could be tuned for both your interest profile and your actual history of interactions/purchases/internet usage, etc. Imagine a world in which you can opt into this tracking, control the data you provide and how it’s used, clear it out and redact it as you please, and opt in and out of responses that are personalized against it. Reasonable people can disagree, but that might actually be useful.
My prediction, though: that’s not gonna happen. OpenAI will first build out the system to collect click and conversion tracking measurements, then they will turn around to advertisers and say “look at how good our conversion rates are,“ and then they’re going to build an explicit ad platform that enshittifies their chat products.
Yeah Google benefits you by tracking you and delivering more relevant ads too. Doesn't make it any less creepy.
The technology isn't anything new but, it's the context that makes it so uncomfortable.
People have very different "expectations" of privacy when they're having a conversation with an AI VS when they're browsing something like Facebook
Not to mention Facebook is free whereas you pay for a GPT subscription
I'm old enough to remember when people had expectations of privacy on Facebook, when the discovery that they didn't made for schocked headlines.
You know, Facebook, built by the guy who scraped college databases for personal information without the student's permission.
There was never any expectation of privacy if you knew Zuck's history. Facemash almost got him expelled for violating individual privacy.
That's the trick: in the early years, most people didn't know Zuckerberg's history.
Caveat emptor, the saying goes. If people don't have an appetite for privacy, don't expect them to take a principled stance based on outrageous headlines.
It's not a coincidence that federally backdoored spyware like Windows is the most popular operating system in the world. I too remember the shocked headlines of the Snowden revelations, and ten-plus years later I work shoulder to shoulder with people that couldn't possibly care less. There is no expectation of privacy, HN too quickly extrapolates it's own virtue signalling to normal people that enjoy using spyware like TikTok, Facebook and Windows.
Back then I didn't have expectations of privacy, I think. I just didn't have a notion of privacy. I didn't think the internet was hostile, I believe?
ChatGPT is also freely available
I don’t think there is any consistency in behavior. People react viscerally to the unproven belief that the Facebook app records conversations. They’ll also happily buy big TVs despite the fact they’re very much listening to you and we have evidence that it’s for ads. Most people have no idea how ads “follow” you, or how companies can figure out what you talked about by connecting the dots from other user metadata.
Was considering buying ads on ChatGPT. A damning thing is that yours audience then are users too cheap to buy a sub...
That's a problem with most unsolicited online advertising. Are YouTube ad-watchers any more likely to splurge on your SaaS?
It's slightly different though, considering a lot of ad spend is for B2B products/services, and a lot of work pay for ChatGPT, so its only people who's work won't pay for a pro subscription.
Whereas, YouTube is for personal consumption in almost all cases, doesn't say anything about your employers willingness to invest in software/services
Just like mobile apps, subs will never come close to the growth expectation. Given trillions are tied up in AI, subs aren't going to do it.
Well, I don't go to website anymore... so ha! ;)
Doesn’t Gemini or whatever Meta’s agent is do this too?
I spent 10 seconds skimming this blog post until realizing it's just AI generated (ironic) slop describing how third party cookie tracking works. https://en.wikipedia.org/wiki/Third-party_cookies#Mechanism
Everyone would be far better off if the author just posted "OpenAI uses third party cookies [wikipedia link]", except maybe the author who wouldn't as many subscriptions for his "threat intel" newsletter.
There are specific details about how the mechanism works like the cookie name (__obi) that can help defenders.
I'm not aware of any filter lists that blocks based on cookie name. It's almost always done at the URL level. Moreover it's bog standard 3rd party cookie tracking. At least for the purposes of "help defenders", there's nothing in the blog post that couldn't be found in 10s with browser devtools.
The audience of people who can read and understand an explanation like this one is probably quite a bit larger than the ones who can replicate the experiment themselves.
And larger still than the ones who would bother to replicate the experiment themselves.
Yeah, I could tell few sentences in that it's Claude-written article. The style has become that distinctive. Hell, a third of the articles I opened on the HN front page today carried that distinctive style.
Doesn't really matter if the content is worth it (and I just realized that recognizing Claude in this also makes me wary of ways this could paper over key details - at this point I start to recognize from my own experience where Claude may be papering over something it didn't actually bother to check).
That just goes back to my original point, which is that this is AI slop and everyone would be better off if OP referenced a canonical description of what's happening, rather than wasting everyone's time by generating 1200 words of AI slop for them to wade through, all for some "threat intelligence" newsletter.
I didn't look at the prose closely enough to check for that, but people wrote inflated explanations of basic things like this all the time pre-LLM. The Wikipedia article doesn't know the specific cookie name, and can't show the results of an experiment verifying that it is in fact being used for ad tracking, or identify specific sites using it in collaboration with OpenAI.
And those got a pass because at least you could defend them with some excuse about how it's some budding author trying to hone their writing skills, or trying to improve their understanding by putting pen to paper. Cases where those excuses don't work (think crappy content marketing pieces from random companies) got short shrift as well. Now for all you know, it's just some dude who prompted claude to "write a blog post about openai's ads".
Time to start having separate container for every domain. I'm pretty sure this should be possible in Firefox.
Firefox (effectively) already does that by default.
https://support.mozilla.org/en-US/kb/introducing-total-cooki...
Doesn't really matter, they can track you in several ways.
Because at least one mechanism of tracking still exists we should just give up on blocking any of them?
SpyGPT.
Is this how random sites know what I search for?
I am once again happy that the EU is fighting practices like these via legislation.
Some outcomes can be annoying, but the net result is still positive, for the consumers and their data privacy at least.
Unfortunately, most LLM companies always were naturally an intelligence operation on civilians, and determining if they are now state-backed seems like a irrelevant detail.
Privacy has been dead for some time now. The fact is most business people never figure out how they are exploited. The modern intelligence campaigns just made it economical to hit almost everyone regardless of scale... often under some silly pretense like terrorists wanting our underpants. =3
I don't know if the destruction of many industries in the EU--including manufacturing--can be called an annoyance.
All the EU bureaucrats can do is regulate, that's why it's so difficult to do business in the EU and that's why there is so little innovation over here. If Microsoft and Apple were started in the EU they would've been shut down within a week because you can't operate from a garage.
I guess that every once in a while this might have a positive outcome for consumers, a broken clock is right twice a day.
I'm sure Microsoft and Apple could have been able to afford a cheap office given their parents' generous financing.
And, regarding the original topic, we wouldn't have many of these privacy issues if the US had not strategically failed to regulate its obnoxious tech monopolies
I don't see how that's relevant? Regulation might have limited some innovation in tech, but gdpr does not affect any of the dying industries. That's more from competition with china, companies with too much interest in giving payouts to shareholders than to innovate, and US financial industry buying most of the best players
I'd classify mandatory encryption backdoors as an industry crisis rather than an annoyance.
Which has nothing to do with the stalking of the adtech industry, unless you are suggesting that the EU might sell access to the keys to the likes of OpenAI?
[though you are not wrong that encryption backdoors are a backwards step on individuals rights to privacy]
As long as you manage to navigate all the dark patterns and not accidentally give your "informed consent".
Malicious compliance (or as it more usually is malicious noncompliance that has not been sufficiently punished so the slimy buggers feel free to continue) is something you should blame the stalkers of adtech for, more than the EU. The legislators could perhaps have made the definitions less wavy, and been harder with enforcement, but that doesn't make it all their fault.
Are you equally happy that the EU is attempting to mandate backdoors into technological platforms and outlaw private encrypted communication?
Copypasting my reply to someone who made almost exactly the same point in a sibling comment earlier:
Which has nothing to do with the stalking of the adtech industry, unless you are suggesting that the EU might sell access to the keys to the likes of OpenAI [though you are not wrong that encryption backdoors are a backwards step on individuals rights to privacy]
Things could be worse. It could be like most of the US where there are both no (or at least far fewer) protections against invasive behaviour of private companies and the government actively trying to backdoor private comms.
crazy dude. chatgpt can't even see the full page content when searching the web and you're saying this...
The ideal personal assistant should know what you do on other websites, and your personal files etc. But it should be local.
Even if I had an actual human person assistant, there would be parts of my life I would want to keep private.
Absolutely not. An ideal personal assistant should know about you exactly what you want them to know about you, and usually that's in a business context only.
Oh, the lifechanging technology isn't completely 100% free? They sell ads just like a bajillion other things?
ChatGPT is not 100% free in the first place, and this is about tracking and selling your data, not the existence of ads themselves.
Just remember it is the same ex-Meta employees who are now at OpenAI adding Ad features into ChatGPT.
Why? Because they are saving humanity. /s
Facebook tracks you outside of their own website for the same reason, and now ChatGPT does the same for the sake of...Ads.
I told you so. [0]
[0] https://news.ycombinator.com/item?id=48996936
Does blocking `bzr.openai.com` on the DNS server prevent this? (without breaking ChatGPT)
Take this with a grain of salt or as a very edgy, polemic take:
I think we really need to hold the individuals running ad networks personally liable for the violation of our privacy rights.
Suspiciously sounds like they need an escape out of the profitability lie.
/tinfoil-hat
why is OpenAI wasting their efforts on implementing ads if they are apparently close to AGI?
I would assume the skill set for implementing ads is quite different than the skill set for developing new models, so it’s quite reasonable to expect different teams are working on each
AGI will most likely have told them that this is the cheapest way to become rich.
AI is really becoming a god that sees/knows everything. Just what we needed ...
Creepy as fuck. I'm really suprised they don't understand quite how user-hostile this is.
Incredibly creepy. I'm amazed they don't have any idea quite how user-hostile this is, or the reception it would receive.
For browsers that implement strict cookie partitioning, the privacy concern is moot. With this, the cookie jar that's used when chatgpt.com is a 3rd party site is completely separate for the jar used on the 1st party site. So your chatgpt.com account can't be linked to your ad views.
Since this has been standard ad tech for awhile, browsers have reacted to this to implement cookie partitioning for exactly these privacy reasons.
The entire ad tech industry may truly be an intelligence agency front. The lengths they go to track individuals is quite alarming.
Much of this information plus a shit ton of other information (ie, LEOs have access to credit reporting) can be bought by governments from the shady data broker networks already.
If people still ignorantly claim this isn’t a George Orwellian dystopia …
Never attribute to malice what can be explained as greed.
The quiet part you rarely hear is that advertising is a smoke and mirror industry akin to throwing darts at a wall. The idea of "homing darts" that stick the target a percentage more of the time would be very attractive in this analogy.
Tech turns advertising from mostly a buckshot spread into something that kinda sounds like something solid and real ("Look, numbers! CTR! CPC! KPI! Our ad product WORKS! Paying customers for your business, guaranteed!")
Ad budgets almost never correlate with actual performance.[0] We're all just lying to ourselves that this business model is performing as expected, and I don't think a reckoning is that far off.
0. https://www.forbes.com/sites/augustinefou/2021/01/02/when-bi...
Highly recommend you use Brave Browser and avoid all this nastiness. If you haven't browsed the web in a browser like Chrome recently, you should try it in a VM or something. It's GRIM. The Web is a wasteland and it's gotten worse because ai chatsites are eating their lunch to the ads are even worse than 10 years ago.
https://brave.com/
https://www.pangram.com/history/c3ad864f-2a50-4785-a0ef-71f2...
why not use your own words? If you are gonna ai generate this blog, just post the prompts instead.
There's a certain irony in linking to someone else and then saying "use your own words". If we follow that, we get a pretty obvious answer: sometimes someone else can express it better and faster.
Most people do not, in fact, want to read raw prompts.
How is this adding to the discussion? The post in itself is quite informative.
Do you constantly want to be reading ai-generated content on this site? If so, why not just stay on chatgpt.com and ask it to generate what hackernews.com would look like today? It's adding to the discussion because I feel like the author broke a social contract by probably putting less effort into writing this than I did reading it.
Why comment if you have nothing good to say?
Love how there's an extremely lengthy explanation of the oldest way to track people on the internet like this shit hasn't been happening since cookies were created.
This might help:
``` ||openai.com^$cookie=__obi ```
Does this look right? AdGuard style.