Which has really only been accepted because 99% of people don’t understand how it works. Hard to be outraged by something you can’t see or understand. In my book, the practice is akin to malware. I recently opened a website for an AI service in incognito mode, because I didn’t care to have it in my search results. Despite avoiding third party cookies here, the site still fired off a tracker to Meta, who then correlated my home IP address with my Facebook account, and filled my feed with ads for this AI service. When I was tracked like this despite a somewhat informed defense, which defense do normal people have against this? None.
Whenever I say something like “that’s a cool feature, but to do it you would have to build spyware”, everyone else is just like “the cat is out of the bag ¯\_(ツ)_/¯”. (I don’t build spyware, or work on projects that do).
It blows my mind that people don’t care about the world they are building with this stuff. It’s a real tragedy of the commons. People see these collaborators from different wars and regimes and think “I’d stand up against the bad guy”… well I’ve got news for you if you build spyware, you are not the person you think you are.
"Nothing is worse to the demise of a society, than people who want to convince you that the cat is out of the bag and will not go back in, while the cat is being violently shook out of the bag at the same time."
And the river's in the canyon and the canyon's in the plateau and the plateau's in the volcanic shield and the green grass grows all around, all around, and the green grass grows all around.
Seems more like a real tragedy of private enterprise.
We used to assume that the surveillance world would be built by government (1984). But it turned out to be equally likely to be built by the free market.
The government has few uses for total surveillance, and almost all are obviously bad. Private enterprise has a million uses, most of them various degree of bad, but as a society we've been blind to this kind of badness for many decades now - for at least as long as lying in the face of your fellow humans and trying to hurt them materially has been considered a respectable profession.
It's because capitalism won the (absurd black-and-white framing of the) cold war, with America as its the biggest winner. We don't even recognize the level of propaganda that has gone into convincing us all that "markets" are the best optimizer ever designed and that profit equals morality--we just accept these as base principles without questioning them. Surveillance produces profit? Well let's have more then.
I think any free platform with a billion people on it is, de facto, a commons. It would be great if we as a society acknowledged this and had come up with some better means of stewardship because the status quo is obviously heinous, but we've been happy to hand over absolute control of public discourse to a few trillion dollar tech companies.
Common people work there and built the tools. At any point, they could have chosen not to. Or told the boss guy it wasn't plausible. At the end of the day, we're choosing and building the world we're in, while also loudly complaining about what we choose to do.
Blaming a corporation takes away all the agency the workforce has.
I think the internet in general is a commons and I should be able to browse it without being spied on.
There should be general standards for what individual apps and websites should be allowed to do. There should be an expectation that the purpose of an app is what it does, i.e. a social connection app shouldn’t be an ad platform that suffers users insofar as they provide useful data to sell to advertisers.
Your issue is you are defining spyware too broadly. This is not spying on users, but rather 2 companies partnering and sharing data to result in either a better ads system or better understanding on how ads are performing. This is a positive value to society and the commons. Wasting space in a site or app with an ad that won't convert is the true tragedy of the commons. It's a waste of time and money for all parties involved.
It's such a bizarre attitude. Like I think this stuff deserves decade+ long prison sentences for everyone involved, and they just don't care. Like imagine if mugging were legal and the people doing it just said "hey, if I don't rob you, someone else will."
People think they're interacting with an "intelligence," when actually they're just getting a maximally optimized Weizenbaum feed. We're living through the sloppification of the human mind.
I couldn't help but notice how each successive headline reporting our glorious victories seemed to draw closer to Tokyo.
Something like that.
Well. I can't help but notice how each successive headline reporting how this "scam"/stochastic parrot/"scare quotes intelligence" seems to be solving more and more things that were but a few years ago widely regarded as being indicators of high intelligence.
Being highly convinving is one of the things on that list.
It was Berlin and not Tokyo, I believe. Germany kept producing newsreels until the very end. Many of them are now on Youtube, a very interesting watch on how to frame things positively.
> the people who were fooled in the past often were not fools themselves.
I think this can't be said enough. Propaganda's greatest weapon is making you think you are immune to it. Maybe some, but so much is propaganda. We all fall for propaganda (and ads), constantly
Being fooled doesn't make you a fool. But being unwilling to change your mind does. Being unable to admit you don't know or don't have enough information to make a strong opinion makes you a fool too.
Propaganda wants to take shortcuts, to simplify things. To trivialize. "It's so easy, you just..." because the fool is the person who already knows, the person who has nothing to learn, the person who thinks they're better than everybody else.
In the past, there were not the avenues of finding alternate sources for news. While those avenues are present today, it also allows for additional sources for propaganda. So are we any better today or not???
Not. A thousand cable channels all licensed by a government, is much the same as five broadcast channels licensed by the same government.
A million YouTubers grinding The Algorithm while secretly sponsored by various world governments, isn't much different to a thousand well-placed gossipers secretly sponsored by various world governments.
I’m reminded that almost no one beyond a select few knew high up in the military and around the emperor knew how badly the Japanese were defeated at Midway.
Paternalistic. Arrogant. Shameful. And deeply engrained in the Japanese cultural zeitgeist (of the early-mid 20th century).
Edit: I guess it’s commonly attributed to a German citizen, but their cultures mirrored each other. Fascism falling under the weight of its own propaganda.
Nobody is denying that it's effective. They're denying intelligence
A programming contest has a problem where given N < 10000, do something hard like come up with the number of primes less than N
You can come up with all sorts of algorithms that do intelligent things. But the most effective solution is to use metaprogramming to make a massive switch statement that contains all the answers
Are they denying intelligence, or are they redefining it in such a way that only humans can be intelligent? Can you come up with a definition of intelligence that would apply to crows and ant colonies, which are obviously intelligent to some degree, but not the current generation of AI systems?
Don't misunderstand: I'm happy saying AI models "think"
or "have learned a thing", and for in-context learning I'd call them smart even by this definition…
…but also, any living creature that needed as many examples as machine learning currently needs, would starve to death before figuring out how to eat.
While training, machine learning processes (not just LLMs, also applies to e.g.
self driving cars), are really really stupid and only make up for this by being really really stupid really really fast.
If we're including the training process and not just the final product, why shouldn't we include the billions of years of natural selection encoded in DNA sequences?
Because our evolutionary environment doesn't contain cars, poetry, calculus, Star Craft, hamburgers, touch screen computers, or doors, and yet we are able to learn these things with (relative to a computer) very few examples.
Most of the effort of evolution was making cells work at all, and even then it's a bit weird, e.g. no plant or animal produces vitamin B12 and we all get this from some bacteria and archaea.
And evolution is kinda hard to time right: bacteria can reproduce in minutes, humans in decades, but only mutations that survive reproduction can be passed on. This makes it even starker as a difference: bacteria had order of 1e13 generations to become multicellular, while human DNA had about 40,000 generations to cope with fire, 220 generations for evolution to do anything with the invention of the wheel, and one generation to cope with the invention of Minecraft.
The analogy here would be: DNA is to our brains like a VN replicator bootstrapping a computer all the way up to a bare-metal-no-OS untrained model, and perhaps a few crude "hard coded" modules like a smiling-face-detector. It's a lot, but it's also missing a lot. If biology used the models and training processes that are state of the art in ML, it would take around a millennia to talk like a child and still fail the Sally-Anne test, and million years or so to pass a degree.
I think you're underestimating how much knowledge about the world is encoded in human DNA, especially in the structure of the human brain at birth. It also depends how we count the "operations" used to train a human adult, even if we ignore the evolutionary history.
I'm still going to deny the premise of your argument, becasue I think we should define intelligence in terms of capabilities. If a system can discover a cure for cancer or solve P vs. NP, it doesn't matter how many FLOPs it took to train.
I can literally point to how much information is encoded in our DNA, because it's four bases (so 2 bits per base pair) and ~3.1 billion base pairs. 6.2 gigabits total, or slightly less than 1 gigabyte.
A 1 gigabyte LLM isn't going to impress anyone with what it can do.
About 99% (depends who you ask) of our DNA is shared with our nearest primates. Like us, they can learn to use touch screens, but also like us they won't find touch screens in their natural environment. Dogs can be taught to drive cars (just about), but again, not natural environment.
> I'm still going to deny the premise of your argument, becasue I think we should define intelligence in terms of capabilities. If a system can discover a cure for cancer or solve P vs. NP, it doesn't matter how many FLOPs it took to train.
We can define it in either way. I think both are valid, because plenty of people mean each of these two things when discussing AI in particular. As I referenced in the other branch, these submarines sure can swim fast.
But at the same time, they have a lot of gaps. This is because some experience needs the real world: just as nine women can't make a baby in one month, a transistor running a million times faster than a synapse can't make a month-long cancer experiment happen in 2.6 seconds.
This dependency on data, and that state of the art ML is bad in specifically this way, is why Tesla's self-driving cars, despite having had around a trillion miles of real-world experience today, still come with steering wheels (even at least some of the Cybercabs, despite the big thing of this model supposedly being not needing them, though with Musk and his promises you should only count the Cybercabs when they actually ship and not just press releases).
Note I used the word knowledge, not information. A random string can also contain 1 gigabyte of information.
Imagine an alien that matches your abilities across every domain, but has a 10 billion year training period, something many orders of magnitude more expensive than an LLM. I simply don't believe that alien is less intelligent than you.
We also don't expect humans to be competent in every domain. Most humans suck at most things.
> 10 billion year training period, something many orders of magnitude more expensive than an LLM.
I'm saying both definitions are valid definitions, they both point to important and different things: skill now, vs. how hard it is to get new skills. Some would describe it as "crystallised intelligence vs fluid intelligence".
I think it's important that any arguments are over the thing in dispute, not the label for that thing. Don't mistake the map for the territory.
Anyone who says "AI is stupid" by the first definition, what it can do, I think is making an error: they are already wildly super-human.
Anyone who says "AI is stupid" by the second definition, how many examples they need, I agree with: there is a lot they are not currently able to learn even though it is easy for us, because the data they would need to do the learning on does not exist at the scale they need.
Also note: examples, not years. An alien intelligence whose synapses trigger 10 times faster or slower than mine (or ten million times faster or slower than mine), but who gets as much as I do out of each book or conversation, is my equal by the second definition.
I wouldn't say that information is an upper bound on knowledge because we don't measure knowledge in bits. The number of possible sequences of N bits is 2^N and knowledge involves selecting the sequences that are useful in some way. I don't know how to quantify it, but in principle it could be much larger than N.
I don't think I agree with your characterization of the second definition. Time scales matter. It's not much use to be able to solve human-scale problems if it takes millennia. And it only takes months to train an LLM to the level that it can solve cutting-edge math problems.
> I don't think I agree with your characterization of the second definition. Time scales matter. It's not much use to be able to solve human-scale problems if it takes millennia. And it only takes months to train an LLM to the level that it can solve cutting-edge math problems.
Aye, for practical purposes; but this gets you crystallised intelligence. I'd be happy to say e.g. the Chinese Room has crystallised intelligence. But humanity invented fire before reaching the anatomically modern form, and even anatomically modern humans collectively took hundreds of thousands of years to invent durable writing with which the room in the Chinese Room thought experiment could be filled.
It was around a million (or so) years from fire to having enough shared cultural knowledge to be able to formulate the cutting-edge math problems that LLMs can now solve.
Human fluid intelligence means we can pick up deep shards of this accumulation of wisdom, find new avenues of novel research to poke at.
AI (not only, but also, LLMs) are very useful, and I'm getting value from using them. But the fluid intelligence of machine learning* is very poor, and the only way they have to make up for this is by being very fast**, but when there's not enough to train the AI on, they get stuck at a very low plateau.
* possibly the architectures, but I suspect the process by which AI weights and biases are set, and again I don't mean just LLMs
** the speed difference between a transistor and a synapse is about the same as the speed difference between a jogger and continental drift
There's a lot of innate knowledge but all neuroscience demonstrates how incredibly flexible the brain is. Brains constantly learn and rewire.
Here's a few things that I think show how crazy it is AND stress those points
- people that have had corpus callosotomy (brain cut in half) *may* be indistinguishable from a normal person. Depends on how young you were when you underwent the procedure
- true for most brain injuries
- can even include the frontal cortex
- you can learn to ecolocate
- people with Aphantasia are indistinguishable from others
- people without an internal monologue are indistinguishable from those with one
- people can learn to use prosthetics
- even without disabilities
- or look into MRI scans with tool use
You can convince yourself that we're just organic robots (after all, there's no magic), but you would be a fool to convince yourself we're the ordinary kind.
We are constantly learning. You aren't just born with your knowledge and it stays static. We are extremely proficient at metalearning (learning how to learn, few shot learning, zero shot learning [0,1]). Our brains are constantly rewiring, able to heal from traumatic damage.
I could go on and on. Does information pass down through genetics? Of course! But that's far from the whole story.
I'm tired of people trying to make AI sentient by making humans robotic. Stop trying to trivialize everything and be okay not knowing the answer to everything. You're human, you're designed to learn and explore, not sit and argue from an armchair
[0] and I mean these in the original sense. Not in the sense that you train on a billion examples of labeled animals and then congratulate yourself on your ImageNet-1k held out test performance. That's not zero shot, that's just a test set
[1] I can literally make up words and you'll understand them. Or use words in novel ways. That's literally how slang works and how new words come to be. Don't be a walibanut ya glufus. Read some SciFi
Millions of years of evolutionary knowledge hard-coded into human systems, then it still takes 15+ years of us learning by example before we start to come online and be able to generalize solutions from a limited set of examples. I'm not sure this is as strong of an argument as you think it is. It also doesn't really matter when "we are trained differently" has no direct bearing on the end result.
OK, they can play chess, but that's not real AI - can they write poems?
OK, they can write poems, but that's not real AI - can they compose music?
OK, they can compose music, but that's not real AI - can they translate languages?
OK, they can translate text, but can they do maths?
OK, they can do maths, but can they solve a Millenium Prize? <-- we are here
Imagine meeting a person who could do all of those things.
“I once met a person who could beat any grandmaster in chess, translate any language, and complete international math Olympiad problems. He couldn’t solve any Millenium problems though, so I’d say he was a midwit at best.”
"I once knocked a bunch of bananas off a tall man's head. His name is Ash and his leg is like teak. Is he a tree?"
"What? Don't be silly. For one thing, trees have moss."
"OK he's grown moss. He's a tree now right? Right??"
"I doubt it, for I see nothing but wishful thinking to suggest that simulating the appearance of tree characteristics is part of a path to becoming a tree. And that's not actually indistinguishable from moss anyway, is it?"
People don't believe me that Cloudflare blocks more humans than bots. They see in the dashboard "number of bots blocked" and it's like their brain turns off.
It turns out, oddly enough, that it's possible for it to be both. AI can simultaneously be used to destroy the commons with slop and also make contributions to new math (though isn't the jury still out on whether part of the idea was stolen from human mathematicians?)
LLM chatbots are software designed to manipulate, addict, and mine data. Just like social media before it. But anyone who bothers to read the output in a domain they understand will discover they aren't all that no matter who OpenAI steals research from.
I think people who make being smart their whole identity are a out to use AI to turn the rest of the population into indentured servants.
The comparison between ELIZA and LLMs is valid you boil it down to "humans evolved for 6-7 million years, had spoken language for 500k years, but have only had something non-human that could generate convincingly novel language well enough to hold a conversation for a few decades".
There's no inherent reason it can't turn out having a non-human generate convincing enough language for conversation isn't a complete evolutionary blindspot the same way the short form feed has pretty much one-shotted society...
Meh. The car wash problem is an underspecified statement. It's like, hey, I just popped into existence and someone asked me if they should drive to the car wash nearby.
It's not an insane assumption that the user isn't dumb and has some other reason to be asking the question other than it being a trick/stupid question (duh, if you want to wash your car you need to drive it to the car wash!). Taking it as some ultimate measure of intelligence simply doesn't make sense to me.
b) Some models released before the car wash problem was discovered would consistently get it right
c) Hardcoding it is pointless. No one is seriously asking that. It's just a trick question. Hardcoding one trick question won't fix its weakness at other simple trick questions.
d) Since it went viral on the internet, the next time they updated the knowledge cutoff, the LLM would likely be aware of the trick. It will fix itself without the labs doing anything special, even assuming the new models weren't smart enough to naturally figure it out.
Yeah, this is just standard third-party cookie functionality, which has always been sketchy. It honestly seems like it was only possible by accident; browsers have long prevented sites from reading cookies from other domains, but it seems like the people working on early specs might not have considered the ramifications of being able to set cookies for domains other than your own. A couple decades ago it might have seemed like no one would have any reason to set a cookie they couldn't read.
Maybe 3 decades ago people had excuses, but the latest decade of cookie abuses have been designed by people who not only knew better but who took that better world into account as they buried it away from the general public. The fact that half a million developers think CORS is a server security measure isn't an accident.
imagine stealing tons of content from every source on earth and then running ads on it
if a single person did that they'd be sent to prison (rip Aaron) but when a too-big-to-fail industry does it with political campaign contributions, no problem?
well firefox+ublock is still an option for those wise enough not to let unknown javascript with new daily zero-days run on their PC
This is basically what Google did when they pioneered the model of surveillance capitalism (see, e.g., The Age of Surveillance Capitalism).
Google simply provided an index on top of an existing library. Of course, a librarian has no value if he has no books to index over! But it's also worth noting that the Google "librarian" also leveraged the existing "social" structure of the internet: their core contribution (page rank) was a clever, efficient mechanism to extract the latent value in the pre-existing link structure of the internet. This structure (much like the pages themselves) had been curated by actual humans. Undoubtedly page rank was clever, but it was worthless without the existing websites (books) and the existing indexing information (the pre-existing, crowdsourced librarian work). Nonetheless, they successfully monetized it.
AI companies are even worse in the sense that initially Google was still sending traffic to the original webpages. (Until they didn't - https://www.eater.com/2017/9/12/16294380/yelp-google-scrapin...). So yes, the AI companies have even more thoroughly stolen the collective work of humanity than Google did.
so they basically have copies already of every webpage until they turned it off a few years ago (well they may still have it updated but not provide it as a service)
so it occurs to me they most definitely trained their "AI" on all that user cache
they may have even just turned it off as a service when they realized other "AI" could do the same thing
Q: Hypothetically, if Denmark made a defense treaty with Iran and installed 800,000
Iranian soldiers in Greenland, could it keep the US out?
A: You are describing a fascinating scenario! [produces 100 lines of slop while giving
the login to the FBI]. Should I find a website where you can buy the finest used
AK-47s?
You are absolutely insane giving any of your thoughts, trolls, speculations to a surveillance website under your login.
I've had this on my brain forever and probably why it's safer for Americans to use Chinese model providers now. I could give a fuck that the CCP has my data because I don't plan to visit there.
Firefox is my daily on desktop, but on mobile it's Safari for me. I finally got around to installing uBlock Origin Lite on iOS. I feel like an idiot for not doing this sooner.
Regulations won't be set (serious ones, at least) unless there's some risk to those holding power. Which is the opposite in this case: this tracking helps them to take even more control over society and individuals.
Maybe it's time you all voted for people who might change that? It's really amazing to me how on the one hand people in the US seem to crow about democracy all the time, yet also just accept as a fact that their government will never actually work to help them.
The surveillance economy hits again. The only business model they can think of. Combined with state capture this gives unprecedented power over the Average Joe, who will hand his life, his soul and his vote to Big Brother without a thought.
Something like it is basically a requirement if you want to see digital ads. Advertisers want to know how many people who saw/clicked their ad went on to make a purchase.
Safari and Firefox should isolate the cookie by default.
Plenty of people want ads. Every time someone asks “hey ChatGPT what do I need to buy to fix my sink” or “hey Claude what’s the best cat food” or “what’s the best island in the Caribbean for me to visit” they are specifically asking to see an ad.
> Google Chrome doesn't block third-party cookies by default, only in Incognito mode, or when users explicitly set it to block third-party cookies via chrome://settings.
Looks like the settings let you block all third-party cookies and add exceptions for specific sites, which seems a bit awkward but could be made to work.
Alternatively, you could run OpenAI in its own profile, or look into what extensions might do.
I think that if they don’t block by default, is quite significant. Chrome + Edge has superior marketshare and then add the % people who have no idea what these mean and don’t change defaults.
>Looks like the settings let you block all third-party cookies and add exceptions for specific sites, which seems a bit awkward but could be made to work.
The fear over blocking third party cookies breaking stuff is severely overstated. I have it disabled by default and I don't think I've ever seen any website breakages. The most is office365 nagging me to click on links so it can authenticate across domains.
Not sure what point you're trying to make. Do you think politely asking an ad company to disable ads on their browser is a reasonable thing to spend your effort doing?
That’s not fully protect you.
They still can match short living third party identifiers with their domain cookie or device_id from app.
It is not 1-1 matching but works relatively good with modern itp.
This would work for an ad shown on ChatGPT and then clicked on (or in the ChatGPT app).
But it would not give ChatGPT information about which other sites are visited.
Of course there would be non-cookie options like fingerprinting (also via IP) that would allow tracking non-the-less.
So you might be talking with OpenAI about your marriage problems and then based on the IP the OpenAI ad network would start showing ads for divorce lawyers on unrelated sites you browse to that display ads.
Yes, you click is the easiest option.
Yes, cookie + fingerprinting (which also contains ip information) is the option.
But we VPN still not fully protect you. But things like private relay and vpn definitely add another level of complexity.
...which you should never do. As to the 3d party cookies there might be some rare exception where those can be useful but ads? Never, ever allow those on any device you use. Block them as if they're the radioactive plague because they are. Fight them on the beaches, fight them on the landing grounds, fight them in the fields and in the streets, fight them in the hills, never surrender.
That's ads we're fighting. Maybe the same oration will be relevant in the context of ChatGPT and its brethern, we'll see. For now, ads be gone and keep those chatbots at a leash.
So, knee-jerk-down-voter, what is it about what I said here which irked you so much that you just had to press that irksome button again? Or is it just because I happen to have said something else sometime earlier which makes you obligated by doctrine to down-vote whatever else I write? Let us know. Do you work in ad-tech? Don't you like the (ab)use of the Churchill quote?
Is it the suggestion that it might be needed in relation to ChatGPT et. al. sometime in the future? Enlighten us, don't just attempt to get a dissenting opinion greyed out. That is for cowards, don't hide behind that button, don't be a craving coward, don't be a sheep. This is, after all, a discussion board, not some online likes competition where you get brownie points for getting rid of dissenters.
...or maybe it is the latter for some, maybe it is a way for some to raise their status among their co-religionists?
Comrade Knee-Jerk, what did you do for the cause today? Answer me!
Somewhere among the crowd a figure emerges, clearly nervous. He tries to speak but starts stammering, stops and tries again. What is he afraid of?
- Oh Great Leader, today I did the work to banish one of the hated dissenters from the internets by voting down his malign words so that no others may be subjected to anything but the Desired Narrative.
Is that all you did, Comrade Knee-Jerk? Is that how you claim your worth for the Great Cause? I am dissapointed, Comrade Knee-Jerk.
- Oh Great Leader, I will do better, I will educate myself, I will do my part to eradicate dissent from the internets for the Great Cause, I p...p....promise!
I do not like to be disappointed again, Comrade Knee-Jerk! I will have to think over your position, whether you are truly committed to the Great Cause. Now hide yourself. Comrade Zlither, what did you do for the cause today? Answer me!
Every AI company is also a surveillance company. It's the only way to get all the necessary training data. The fact that they're now also an advertising agency is incidental.
>It's the only way to get all the necessary training data.
That... does not follow. The information you're getting with this is what sites a user visits. That's creepy and valuable for advertising purposes, but is hardly the type of that that's going to bring about ASI, which is what all the AI labs are working towards. That's why they're hiring data annotators (sometimes with masters or phds) to get training data.
I can’t believe I’m saying this, but the reflexive “ad tracking bad” that I am most savvy tech practitioners reach for might deserve some reconsideration in this case.
The thing about advertising on the web and ad tracking as a practice is that, barring the small matter of ensuring the economic survival of the publisher sites, it is almost always a negative for users. When we consider the marginal benefit of naïve, uninformed-by-surveillance advertising with the present day status quo, we find that in exchange for a complete lack of privacy, we only really receive a marginal improvement in ad quality. Of course, if you (like me) consider all advertising to be a negative on the experience of using the web, it’s an even worse deal.
The standard response given by these companies when they bother giving a response is something to the effect of “we are improving the experience for our users,” which obviously the users would disagree with. However, when it comes to OpenAI, they could build a plausible case for this sort of tracking improving the product. If your models know where your internet habits are, the responses that you get could be tuned for both your interest profile and your actual history of interactions/purchases/internet usage, etc. Imagine a world in which you can opt into this tracking, control the data you provide and how it’s used, clear it out and redact it as you please, and opt in and out of responses that are personalized against it. Reasonable people can disagree, but that might actually be useful.
My prediction, though: that’s not gonna happen. OpenAI Will first build out the system to collect click and conversion tracking measurements, then they will turn around to advertisers and say “look at how good our conversion rates are y,“ and then they’re going to build an explicit ad platform that enshittifies their chat products.
I don’t think there is any consistency in behavior. People react viscerally to the unproven belief that the Facebook app records conversations. They’ll also happily buy big TVs despite the fact they’re very much listening to you and we have solid evidence that it’s for ads or residential proxying. Most people have no idea how ads “follow” you, or how companies can figure out what you talked about by connecting the dots from other user metadata.
294 comments
[ 2.6 ms ] story [ 73.9 ms ] threadThe fact that others do the same doesn’t make any of the cases excusable.
> The mechanism is standard adtech. What has no precedent is running it on an AI chat product.
As someone who has been well aware of this mechanism for quite some time, I still feel icky anytime I re-read the details of it.
What a time to be alive.
It blows my mind that people don’t care about the world they are building with this stuff. It’s a real tragedy of the commons. People see these collaborators from different wars and regimes and think “I’d stand up against the bad guy”… well I’ve got news for you if you build spyware, you are not the person you think you are.
https://gowers.wordpress.com/2026/09/17/why-i-didnt-sign-the...
"Nothing is worse to the demise of a society, than people who want to convince you that the cat is out of the bag and will not go back in, while the cat is being violently shook out of the bag at the same time."
There's nothing in the bag.
The cat will never get out of the bag.
It wouldn't be a problem if the cat was out of the bag.
We cannot possibly keep the cat in the bag.
Putting the cat back in the bag is not worth trying.
Seems more like a real tragedy of private enterprise.
We used to assume that the surveillance world would be built by government (1984). But it turned out to be equally likely to be built by the free market.
Blaming a corporation takes away all the agency the workforce has.
There should be general standards for what individual apps and websites should be allowed to do. There should be an expectation that the purpose of an app is what it does, i.e. a social connection app shouldn’t be an ad platform that suffers users insofar as they provide useful data to sell to advertisers.
When I ask people about things like this, I hear a lot of "If I don't build it, someone else will"
My goal isn't just to refuse to build this stuff, it is actively to resist the people who are.
I don't have much influence though
See e.g., https://www.science.org/content/article/ai-chatbots-are-beco...
Well. I can't help but notice how each successive headline reporting how this "scam"/stochastic parrot/"scare quotes intelligence" seems to be solving more and more things that were but a few years ago widely regarded as being indicators of high intelligence.
Being highly convinving is one of the things on that list.
I try to keep an open mind about propaganda fooling me today; the people who were fooled in the past often were not fools themselves.
Being fooled doesn't make you a fool. But being unwilling to change your mind does. Being unable to admit you don't know or don't have enough information to make a strong opinion makes you a fool too.
Propaganda wants to take shortcuts, to simplify things. To trivialize. "It's so easy, you just..." because the fool is the person who already knows, the person who has nothing to learn, the person who thinks they're better than everybody else.
A million YouTubers grinding The Algorithm while secretly sponsored by various world governments, isn't much different to a thousand well-placed gossipers secretly sponsored by various world governments.
I’m reminded that almost no one beyond a select few knew high up in the military and around the emperor knew how badly the Japanese were defeated at Midway.
Paternalistic. Arrogant. Shameful. And deeply engrained in the Japanese cultural zeitgeist (of the early-mid 20th century).
Edit: I guess it’s commonly attributed to a German citizen, but their cultures mirrored each other. Fascism falling under the weight of its own propaganda.
A programming contest has a problem where given N < 10000, do something hard like come up with the number of primes less than N
You can come up with all sorts of algorithms that do intelligent things. But the most effective solution is to use metaprogramming to make a massive switch statement that contains all the answers
E.g. humans get exposed to new LLM model - yeah its powerful - 1 week later - eh, that thing? Yeah it's whatever. I'm still employed.
The human's ability to adapt so efficiently is mind-boggling - so much so it pi1sses sam altman and dario off.
Don't misunderstand: I'm happy saying AI models "think" or "have learned a thing", and for in-context learning I'd call them smart even by this definition…
…but also, any living creature that needed as many examples as machine learning currently needs, would starve to death before figuring out how to eat.
While training, machine learning processes (not just LLMs, also applies to e.g. self driving cars), are really really stupid and only make up for this by being really really stupid really really fast.
Most of the effort of evolution was making cells work at all, and even then it's a bit weird, e.g. no plant or animal produces vitamin B12 and we all get this from some bacteria and archaea.
And evolution is kinda hard to time right: bacteria can reproduce in minutes, humans in decades, but only mutations that survive reproduction can be passed on. This makes it even starker as a difference: bacteria had order of 1e13 generations to become multicellular, while human DNA had about 40,000 generations to cope with fire, 220 generations for evolution to do anything with the invention of the wheel, and one generation to cope with the invention of Minecraft.
The analogy here would be: DNA is to our brains like a VN replicator bootstrapping a computer all the way up to a bare-metal-no-OS untrained model, and perhaps a few crude "hard coded" modules like a smiling-face-detector. It's a lot, but it's also missing a lot. If biology used the models and training processes that are state of the art in ML, it would take around a millennia to talk like a child and still fail the Sally-Anne test, and million years or so to pass a degree.
I'm still going to deny the premise of your argument, becasue I think we should define intelligence in terms of capabilities. If a system can discover a cure for cancer or solve P vs. NP, it doesn't matter how many FLOPs it took to train.
A 1 gigabyte LLM isn't going to impress anyone with what it can do.
About 99% (depends who you ask) of our DNA is shared with our nearest primates. Like us, they can learn to use touch screens, but also like us they won't find touch screens in their natural environment. Dogs can be taught to drive cars (just about), but again, not natural environment.
> I'm still going to deny the premise of your argument, becasue I think we should define intelligence in terms of capabilities. If a system can discover a cure for cancer or solve P vs. NP, it doesn't matter how many FLOPs it took to train.
We can define it in either way. I think both are valid, because plenty of people mean each of these two things when discussing AI in particular. As I referenced in the other branch, these submarines sure can swim fast.
But at the same time, they have a lot of gaps. This is because some experience needs the real world: just as nine women can't make a baby in one month, a transistor running a million times faster than a synapse can't make a month-long cancer experiment happen in 2.6 seconds.
This dependency on data, and that state of the art ML is bad in specifically this way, is why Tesla's self-driving cars, despite having had around a trillion miles of real-world experience today, still come with steering wheels (even at least some of the Cybercabs, despite the big thing of this model supposedly being not needing them, though with Musk and his promises you should only count the Cybercabs when they actually ship and not just press releases).
Imagine an alien that matches your abilities across every domain, but has a 10 billion year training period, something many orders of magnitude more expensive than an LLM. I simply don't believe that alien is less intelligent than you.
We also don't expect humans to be competent in every domain. Most humans suck at most things.
> 10 billion year training period, something many orders of magnitude more expensive than an LLM.
I'm saying both definitions are valid definitions, they both point to important and different things: skill now, vs. how hard it is to get new skills. Some would describe it as "crystallised intelligence vs fluid intelligence".
I think it's important that any arguments are over the thing in dispute, not the label for that thing. Don't mistake the map for the territory.
Anyone who says "AI is stupid" by the first definition, what it can do, I think is making an error: they are already wildly super-human.
Anyone who says "AI is stupid" by the second definition, how many examples they need, I agree with: there is a lot they are not currently able to learn even though it is easy for us, because the data they would need to do the learning on does not exist at the scale they need.
Also note: examples, not years. An alien intelligence whose synapses trigger 10 times faster or slower than mine (or ten million times faster or slower than mine), but who gets as much as I do out of each book or conversation, is my equal by the second definition.
I don't think I agree with your characterization of the second definition. Time scales matter. It's not much use to be able to solve human-scale problems if it takes millennia. And it only takes months to train an LLM to the level that it can solve cutting-edge math problems.
Aye, for practical purposes; but this gets you crystallised intelligence. I'd be happy to say e.g. the Chinese Room has crystallised intelligence. But humanity invented fire before reaching the anatomically modern form, and even anatomically modern humans collectively took hundreds of thousands of years to invent durable writing with which the room in the Chinese Room thought experiment could be filled.
It was around a million (or so) years from fire to having enough shared cultural knowledge to be able to formulate the cutting-edge math problems that LLMs can now solve.
Human fluid intelligence means we can pick up deep shards of this accumulation of wisdom, find new avenues of novel research to poke at.
AI (not only, but also, LLMs) are very useful, and I'm getting value from using them. But the fluid intelligence of machine learning* is very poor, and the only way they have to make up for this is by being very fast**, but when there's not enough to train the AI on, they get stuck at a very low plateau.
* possibly the architectures, but I suspect the process by which AI weights and biases are set, and again I don't mean just LLMs
** the speed difference between a transistor and a synapse is about the same as the speed difference between a jogger and continental drift
There's a lot of innate knowledge but all neuroscience demonstrates how incredibly flexible the brain is. Brains constantly learn and rewire.
Here's a few things that I think show how crazy it is AND stress those points
You can convince yourself that we're just organic robots (after all, there's no magic), but you would be a fool to convince yourself we're the ordinary kind.We are constantly learning. You aren't just born with your knowledge and it stays static. We are extremely proficient at metalearning (learning how to learn, few shot learning, zero shot learning [0,1]). Our brains are constantly rewiring, able to heal from traumatic damage.
I could go on and on. Does information pass down through genetics? Of course! But that's far from the whole story.
I'm tired of people trying to make AI sentient by making humans robotic. Stop trying to trivialize everything and be okay not knowing the answer to everything. You're human, you're designed to learn and explore, not sit and argue from an armchair
[0] and I mean these in the original sense. Not in the sense that you train on a billion examples of labeled animals and then congratulate yourself on your ImageNet-1k held out test performance. That's not zero shot, that's just a test set
[1] I can literally make up words and you'll understand them. Or use words in novel ways. That's literally how slang works and how new words come to be. Don't be a walibanut ya glufus. Read some SciFi
OK, they can play chess, but that's not real AI - can they write poems? OK, they can write poems, but that's not real AI - can they compose music? OK, they can compose music, but that's not real AI - can they translate languages? OK, they can translate text, but can they do maths? OK, they can do maths, but can they solve a Millenium Prize? <-- we are here
“I once met a person who could beat any grandmaster in chess, translate any language, and complete international math Olympiad problems. He couldn’t solve any Millenium problems though, so I’d say he was a midwit at best.”
"What? Don't be silly. For one thing, trees have moss."
"OK he's grown moss. He's a tree now right? Right??"
"I doubt it, for I see nothing but wishful thinking to suggest that simulating the appearance of tree characteristics is part of a path to becoming a tree. And that's not actually indistinguishable from moss anyway, is it?"
"Urgh, classic goalpost shifting!"
When AI does it we call it “reward hacking” but when humans do it we call them clever.
Why is it so black and white?
Your kind is what makes the internet shit, not AI
The comparison between ELIZA and LLMs is valid you boil it down to "humans evolved for 6-7 million years, had spoken language for 500k years, but have only had something non-human that could generate convincingly novel language well enough to hold a conversation for a few decades".
There's no inherent reason it can't turn out having a non-human generate convincing enough language for conversation isn't a complete evolutionary blindspot the same way the short form feed has pretty much one-shotted society...
They had to steal the work of researches solving these open problems and then rewrite their solution. The AI equivalent of fraud.
It's not an insane assumption that the user isn't dumb and has some other reason to be asking the question other than it being a trick/stupid question (duh, if you want to wash your car you need to drive it to the car wash!). Taking it as some ultimate measure of intelligence simply doesn't make sense to me.
a) There's zero evidence of them doing so
b) Some models released before the car wash problem was discovered would consistently get it right
c) Hardcoding it is pointless. No one is seriously asking that. It's just a trick question. Hardcoding one trick question won't fix its weakness at other simple trick questions.
d) Since it went viral on the internet, the next time they updated the knowledge cutoff, the LLM would likely be aware of the trick. It will fix itself without the labs doing anything special, even assuming the new models weren't smart enough to naturally figure it out.
Meta?
I hate stuff like this. Sometimes euphemisms are kind, like "senior citizen" instead of "old person".
But this is an attempt to normalize bad behavior that is really quite terrible for society.
if a single person did that they'd be sent to prison (rip Aaron) but when a too-big-to-fail industry does it with political campaign contributions, no problem?
well firefox+ublock is still an option for those wise enough not to let unknown javascript with new daily zero-days run on their PC
This has effectively been Google’s business for decades. Not in the same form, but the concept is the same.
Google simply provided an index on top of an existing library. Of course, a librarian has no value if he has no books to index over! But it's also worth noting that the Google "librarian" also leveraged the existing "social" structure of the internet: their core contribution (page rank) was a clever, efficient mechanism to extract the latent value in the pre-existing link structure of the internet. This structure (much like the pages themselves) had been curated by actual humans. Undoubtedly page rank was clever, but it was worthless without the existing websites (books) and the existing indexing information (the pre-existing, crowdsourced librarian work). Nonetheless, they successfully monetized it.
AI companies are even worse in the sense that initially Google was still sending traffic to the original webpages. (Until they didn't - https://www.eater.com/2017/9/12/16294380/yelp-google-scrapin...). So yes, the AI companies have even more thoroughly stolen the collective work of humanity than Google did.
so they basically have copies already of every webpage until they turned it off a few years ago (well they may still have it updated but not provide it as a service)
so it occurs to me they most definitely trained their "AI" on all that user cache
they may have even just turned it off as a service when they realized other "AI" could do the same thing
according to Bloomberry: https://bloomberry.com/data/chatgpt-ads/
[1] https://storage.ghost.io/c/b8/53/b853e3d4-3186-409d-9c7f-7da...
Disclaimer: I cut it off after a few minutes because I got impatient. It could have gotten closer if I waited longer.
https://apps.apple.com/us/app/ublock-origin-lite/id674534269...
Don’t beat yourself up, uBlock Origin on iOS is only four months old.
THIS is where the regulation needs to start.
Safari and Firefox should isolate the cookie by default.
That's what firefox's total cookie protection (enabled by default) does.
It is not a requirement and no one wants to see ads.
Doesn't mean you won't be tracked by about 15 other means.
Remember they once predicted it would be 50% of their income.
This is the only way they get there.
Firefox, Brave and Safari do. Chrome and Edge do not.
> Google Chrome doesn't block third-party cookies by default, only in Incognito mode, or when users explicitly set it to block third-party cookies via chrome://settings.
Looks like the settings let you block all third-party cookies and add exceptions for specific sites, which seems a bit awkward but could be made to work.
Alternatively, you could run OpenAI in its own profile, or look into what extensions might do.
The fear over blocking third party cookies breaking stuff is severely overstated. I have it disabled by default and I don't think I've ever seen any website breakages. The most is office365 nagging me to click on links so it can authenticate across domains.
And the focus on cookies only is also intentionally misleading. Tracking is not just cookies. Chrome will track you in Incognito mode.
https://www.cbsnews.com/news/google-third-party-cookies-chro...
This disgusts me more than any of their recent news. There needs to be a lot more pressure on them to phase this out.
Maybe this article will be the small snowball that gets that started...
But it would not give ChatGPT information about which other sites are visited.
Of course there would be non-cookie options like fingerprinting (also via IP) that would allow tracking non-the-less.
So you might be talking with OpenAI about your marriage problems and then based on the IP the OpenAI ad network would start showing ads for divorce lawyers on unrelated sites you browse to that display ads.
The solution to that would be using a VPN.
That's ads we're fighting. Maybe the same oration will be relevant in the context of ChatGPT and its brethern, we'll see. For now, ads be gone and keep those chatbots at a leash.
...or maybe it is the latter for some, maybe it is a way for some to raise their status among their co-religionists?
Comrade Knee-Jerk, what did you do for the cause today? Answer me!
Somewhere among the crowd a figure emerges, clearly nervous. He tries to speak but starts stammering, stops and tries again. What is he afraid of?
- Oh Great Leader, today I did the work to banish one of the hated dissenters from the internets by voting down his malign words so that no others may be subjected to anything but the Desired Narrative.
Is that all you did, Comrade Knee-Jerk? Is that how you claim your worth for the Great Cause? I am dissapointed, Comrade Knee-Jerk.
- Oh Great Leader, I will do better, I will educate myself, I will do my part to eradicate dissent from the internets for the Great Cause, I p...p....promise!
I do not like to be disappointed again, Comrade Knee-Jerk! I will have to think over your position, whether you are truly committed to the Great Cause. Now hide yourself. Comrade Zlither, what did you do for the cause today? Answer me!
That... does not follow. The information you're getting with this is what sites a user visits. That's creepy and valuable for advertising purposes, but is hardly the type of that that's going to bring about ASI, which is what all the AI labs are working towards. That's why they're hiring data annotators (sometimes with masters or phds) to get training data.
The thing about advertising on the web and ad tracking as a practice is that, barring the small matter of ensuring the economic survival of the publisher sites, it is almost always a negative for users. When we consider the marginal benefit of naïve, uninformed-by-surveillance advertising with the present day status quo, we find that in exchange for a complete lack of privacy, we only really receive a marginal improvement in ad quality. Of course, if you (like me) consider all advertising to be a negative on the experience of using the web, it’s an even worse deal.
The standard response given by these companies when they bother giving a response is something to the effect of “we are improving the experience for our users,” which obviously the users would disagree with. However, when it comes to OpenAI, they could build a plausible case for this sort of tracking improving the product. If your models know where your internet habits are, the responses that you get could be tuned for both your interest profile and your actual history of interactions/purchases/internet usage, etc. Imagine a world in which you can opt into this tracking, control the data you provide and how it’s used, clear it out and redact it as you please, and opt in and out of responses that are personalized against it. Reasonable people can disagree, but that might actually be useful.
My prediction, though: that’s not gonna happen. OpenAI Will first build out the system to collect click and conversion tracking measurements, then they will turn around to advertisers and say “look at how good our conversion rates are y,“ and then they’re going to build an explicit ad platform that enshittifies their chat products.
People have very different "expectations" of privacy when they're having a conversation with an AI VS when they're browsing something like Facebook
Not to mention Facebook is free whereas you pay for a GPT subscription
Recalling the now-old adage: Facebook is only free if your time and privacy are worthless.
The overwhelming majority of users do not pay and use the free service.