"multiple providers disclose sensitive conversation-derived artifacts — including titles, prompts, and screenshots — to third parties, often alongside persistent user identifiers that enable user attribution. We also find that some providers publicly expose conversation permalinks without access controls, allowing trackers to read the entire conversation."
Security by obscurity -- such an age old anti-pattern!
I believe many AI tools like Gemini generate publicly accessible URLs when we click "Share" on any chat conversation -- and expect users to then own the lifecycle of that link
Depending on how the link gets handled -- by the browser, device OS, any hooks/plugins/extensions, aggressive telemetry, social media url previews, preload/prefetch, wrapping and url shortening, etc as it reaches the intended user -- there are countless ways in which the URL can be indexed and scraped
There was a issue not long ago when Claude artifacts were indexed en-masse by Google and other search engines
This is shockingly lax approach to data security and privacy by design
The same assumptions are true about giving any human that shareable link. They could pass it on to anyone, screenshot it, paste it into their own session. This has been true since before "share with link" permissions on Docs and elsewhere.
If you click "provide a shareable link" you should decide (and behave) as though that made it public.
I'm not saying it's good privacy posture on the side of the companies, but how else do you think that would work if there isn't any authentication step for the person viewing it? Even with authentication, "three may keep a secret, if two of them are dead."
Depends on the uuid, if it's long enough and genuinely random then sure, if they follow a predictable pattern and someone could mine shared URLs by trying lots of URLs in some finite range that absolutely violates my expectations of "only someone with this link can view"
not as bad since the blast radius is only your chat vs. knowing your password exposes all your data.
however - agree that this is not great - espeically if chat TTL is long. someone who gets your URL can read everything you're asking (eg. by sniffing your network/accessing your browser history)
I'm shocked... SHOCKED! that these companies would sell this data to advertisers and violate the privacy of their users. Who could have seen this coming??
This is a bit surprising to me, considering how much AI companies love to hoard data. Especially since some of these ad companies are direct competitors!
My best guess is that these ad mechanisms are a bit rushed and/or that investor demands for profitability are fighting against company self-interest.
I'm taking an onboarding process for an Ai company helping other (much) bigger Ai companies and the amount of vibecoded platforms and documentation is staggering. Like, training process so broken that the platform just doesn't load sometimes, things that have never even been tried are published and you have to use them and they suck so much because it is apparent that no human ever touched that and probably wouldn't want to
It's a part time job hiring contractors, I'm doing this only because they pay would be nice but I know very well I can be fired anytime or when the project is done...
Anyway I'm a student so I'm not looking for something stable. BTW now it's getting better but if I were the client paying for those services I'd be pretty fucking furious for what's going on - guess they'll never know tho
So I only read the abstract, but the question is if the data is leaked, by accident, or if it's deliberately provided. My guess is that we're talking about the first scenario, and that this is an accident.
If that's the case, then I'm not surprised at all. Actually I also wouldn't be surprised if they sold the data, but that's a different story. If we look at OpenAI for instance, they have on multiple occasion shown that they do not have the operational experience or resources to run their services in a safe and secure manor, nor do they frankly have an impressive up reliability (in terms of operational stability).
I'd support your guess that all of this is rushed in an attempt to push for profitabilitet/growth.
Well, some data is deliberately provided at least. It's not a situation of other apps managing to grab the data like the facebook-localhost exploit
>The most prevalent third-party services included in CSP headers belong to Google Tag Manager (googletagmanager.com), Google Analytics (google-analytics.com), and Google Ads (googleadservices.com and doubleclick.net). Yet, as Table 8 shows, CSP policies commonly include other prominent actors in the advertising industry, such as TikTok and Meta.
the question is if the data is leaked, by accident, or if it's deliberately provided. My guess is that we're talking about the first scenario, and that this is an accident.
Read the T&Cs. If there's even the tiniest bit of "we might provide your data to third parties for the purposes of...", it's not an accident. Virtually all the AI company T&Cs I've looked at had weasel words that open that door, because it's obvious to anyone paying attention that cramming ads into AI products is the next frontier in AI revenue streams.
I don’t think it matters much if it is intentional or not, no? In both cases it’s a breach of privacy that should be punished. Though I myself do not believe they are selling, I don’t think they reached that level of sophistication yet, they seem to be way too hacky for that type of scheme. But that will for sure come later on
I'd be kind of surprised if OpenAI were really doing this deliberately because a whole bunch of their execs come from Meta, and all those guys learned the hard way.
First: You don't want to leak information about your users to advertising networks because it's going to leak, get back to your customers, they're going to figure out you're doing it and get really angry.
But second and more importantly - it's a much better business model to collect that data for yourself, keep it in house and then you control how you use that data to target ads which gives you a massive competitive advantage in selling ads because you have unique targeting data.
The way meta does this now is the model, they don't give the advertiser a list of the people you're going to show the advert to, the advertiser gives you a list of characteristics they want to hit and meta decides who those people are.
> a whole bunch of their execs come from Meta, and all those guys learned the hard way.
Apparently they have not learned, or they learned the wrong lesson. I do not think they leak it intentionally as hoarding is typically more profitable than selling but anyone who has been in the industry for some time knows that the move fast and break things attitude has caused enormous amounts of data leaks.
> First: You don't want to leak information about your users to advertising networks because it's going to leak, get back to your customers, they're going to figure out you're doing it and get really angry.
And then they will keep using your digital drugs like nothing happened and forget about the whole thing. So watch out, executive!
> But second and more importantly - it's a much better business model to collect that data for yourself, keep it in house and then you control how you use that data to target ads which gives you a massive competitive advantage in selling ads because you have unique targeting data.
Perhaps, but getting to the saturated market on this might be something LLM labs simply don’t have time for. They are haemorrhaging money, no path to profitability and OpenAI especially has made ridiculous promises on data centre -spending for the coming year. They need money now.
> I'd be kind of surprised if OpenAI were really doing this deliberately because a whole bunch of their execs come from Meta, and all those guys learned the hard way.
The lesson they've learned is that if the profit from an activity is X, and the sanctions/reputational hit costs less than X, it's full steam forward.
Thank you. I keep saying this and the spin doctors keep using words that imply "oh, my...I am so sorry we had that tiny little privacy issue...we'll get right on that in the next sprint...". And winning the message war.
It's not a 'leak'. It's in the T&Cs you agreed to. This is the business plan.
I’m not sure. It could definitely be incompetence. We are talking about an industry rushing as fast as possible in every directions, their systems are very likely a complete mess. I really wouldn’t be surprised if advertisers are getting that data for free just because the AI labs didn’t bother
It's because of people like you that these companies have a free ride. When we're talking about billion dollar companies there's no such thing as a "mistake", because they need to be held accountable for everything that they do. Thinking otherwise let them keep the rewards while getting away with anything wrong they cause to society.
I think most people don't realize that there's always organizational and systemic intent in the creation of incompetence. an organization should be accountable for it's practices and worker adherence to policies - it's why safety regulations and training exist, after all. when individuals fail within that system, that's partially the individual's blame but it's also the problem of the institution - did it provide enough training? did it push projects to go too fast? is the core ethos reified enough for everyone to act accordingly?
in this sense, any individualized incompetence is functionally malice-by-omission by organizations, especially ones as large and well-funded as LLM model providers
Those companies are grossly incompetent to the point of committing what is likely to be felonies. If you think I’m deflecting you’re misguided, I’ve been advocating for months for the AI labs to be shutdown and their leadership investigated for gross negligence and potential crimes.
I would still bet they aren’t making any money from data leaking to advertisers
In an old Simpsons episode Lisa gets to visit the Teachers room, where all the staff are making fun of the children. Groundskeeper Willie is pantomiming Milhouse “Oh I am Milhouse, I tell all my secrets to Willie since I have no friends!” and the teachers laugh. Later something embarrassing happens to Milhouse and he immediately runs away crying “I have to tell this to Willie!”.
Inequality has eroded trust in society in ~50 years, a lot of the old models (heh) of how we see the world aren't relevant. It's hard to exist when everything around us is adversarial.
I love how some people think the scale of a problem doesnt matter. Like yeah we has data brokerw 50 years ago, way less of them and the scope they could collect data about was significantly more limited. That matters a lot
It's not inequality. It's who we've chosen to reward. Adversarial people have eaten a lot of the world and that's because we let them.
People love inequality when it's a celebrity they adore living large. Someone who hasn't scammed them and has demonstrably improved their life and the lives of others in a tangible way.
Yes, the media has conditioned most of us to believe that anyone starting or investing in a tech company and at least a billion dollars in equity is "doing good" and should be emulated.
It’s just inequality full stop. Trust is a word by wannabe-Bernays polsci people who would just so dearly want the proles to trust their betters, but for some dastardly material reasons they don’t.
But we are not really like Milhouse. In reality the teachers didn't care about us enough to make fun of us in the teachers lounge. We were just that year's batch of work soon to be forgotten when they move on to the next. This is exactly the same as the advertising companies. You are just a member of the cohort of a particular target campaign. Nobody cares about you or even knows you are part of that target cohort.
Unless you become a target of the government. Then people a lot worse than any teacher you ever had will be looking through it, and they, unfortunately, do care.
I would point you to the insurance companies deciding on premiums based on people’s Amazon purchases as an indicator of risk. There’s plenty of potential for long-term abuse of your profiling data. Nobody cares about you in a specific way, but the algorithms can be designed to target people “like you” and hurt YOU in a very specific manner.
A lot of people seem to be very in denial about the fact that OpenAI and co do not give a crap about you. They don't care about the agreements you've signed. You're just a pile of cash to them
You are currently a cost to them¹ - using your data as free training/refinement, and potentially selling², it is a way to offset that a little.
----
[1] https://isaiprofitable.com/ - some of the green bars are creeping up a bit, but not much unless you count the shovel sellers
[2] sorry, leaking³
[3] Though it could of course be incompetence rather than malice, a mistake they are not actually making anything out of, as per Grey's amendment⁴ to Hanlon's razor.
[4] Any sufficiently advanced incompetence is indistinguishable from malice.
I have recently noticed that e.g. ChatGPT, when used from a web browser, periodically sends unfinished prompts to their servers, namely to the `conversation/prepare` endpoint, without waiting for the user to actually send it.
This partial prompt data might potentially be used to "pre-warm" some kind of cache.
But it may also be used to track the user's writing cadence, error correction style and evolution of their stub ideas as they are being formulated into a prompt. I would assume that such data could also be sold to the advertisers.
This is my concern as well- the same wiggle-room methodology that allowed a business to claim not to “sell or share” PII, because “user data collaboration” was not part of the legal definition prior to CCPA.
OpenAI’s statements in response to the Millenium Prize (and related) disputes I think are a pretty obvious example of this in practice. One man’s “user prompts” is not another’s “reasoning trace scratchpad”.
This comment by Falserum on the mathematics research post articulates it well:
> But it may also be used to track the user's writing cadence, error correction style
I'm pretty sure it is used for this; but rather than for anything nefarious, my guess is that this info is then fed to a classifier model to ensure that users of ChatGPT-the-service (as opposed to the OpenAI inference API) are actual humans, rather than agents trying to circumvent having to pay API pricing.
“Oh, we’ll store your unfinished thoughts, privately typed into a text box, on the off-chance that you want to continue writing it on another device. And no, we won’t ask you nor give you any semblance of true privacy.”
Ad-tech spent 20 years trying to infer intent from clickstreams. Chat apps now hand over an AI-written one-line summary of intent, labeled and keyed to a cookie. Are we sure the ads business model in AI is about ads in the chat, and not the chat as the targeting signal?
It's the same lesson as the Navier-Stokes credit fight earlier this month. Buckmaster and Alpoge had their unpublished drafts in private Codex sessions and OpenAI says nobody saw them but admits de-identified product data may have improved its models. There it's training data, here it's ad trackers. Either way, prompts and results that should stay private don't.
Thats why even though open models aren't perfect it has to win. You can skip the app and run the model yourself.
131 comments
[ 3.2 ms ] story [ 56.9 ms ] threadNot good at all.
Perplexity does this. Visiting a past perplexity search url exposes your full conversation.
It’s not like someone’s gonna guess that URL… right?
I believe many AI tools like Gemini generate publicly accessible URLs when we click "Share" on any chat conversation -- and expect users to then own the lifecycle of that link
Depending on how the link gets handled -- by the browser, device OS, any hooks/plugins/extensions, aggressive telemetry, social media url previews, preload/prefetch, wrapping and url shortening, etc as it reaches the intended user -- there are countless ways in which the URL can be indexed and scraped
There was a issue not long ago when Claude artifacts were indexed en-masse by Google and other search engines
This is shockingly lax approach to data security and privacy by design
If you click "provide a shareable link" you should decide (and behave) as though that made it public.
I'm not saying it's good privacy posture on the side of the companies, but how else do you think that would work if there isn't any authentication step for the person viewing it? Even with authentication, "three may keep a secret, if two of them are dead."
however - agree that this is not great - espeically if chat TTL is long. someone who gets your URL can read everything you're asking (eg. by sniffing your network/accessing your browser history)
The amount of sensitive information that squeaks by on screenshots is pretty large. This is a pretty massive security issue
Someone should have to investigate, but I suppose it's all "legal"?
In EU law it is very likely against GDPR.
My best guess is that these ad mechanisms are a bit rushed and/or that investor demands for profitability are fighting against company self-interest.
Anyway I'm a student so I'm not looking for something stable. BTW now it's getting better but if I were the client paying for those services I'd be pretty fucking furious for what's going on - guess they'll never know tho
If that's the case, then I'm not surprised at all. Actually I also wouldn't be surprised if they sold the data, but that's a different story. If we look at OpenAI for instance, they have on multiple occasion shown that they do not have the operational experience or resources to run their services in a safe and secure manor, nor do they frankly have an impressive up reliability (in terms of operational stability).
I'd support your guess that all of this is rushed in an attempt to push for profitabilitet/growth.
>The most prevalent third-party services included in CSP headers belong to Google Tag Manager (googletagmanager.com), Google Analytics (google-analytics.com), and Google Ads (googleadservices.com and doubleclick.net). Yet, as Table 8 shows, CSP policies commonly include other prominent actors in the advertising industry, such as TikTok and Meta.
Read the T&Cs. If there's even the tiniest bit of "we might provide your data to third parties for the purposes of...", it's not an accident. Virtually all the AI company T&Cs I've looked at had weasel words that open that door, because it's obvious to anyone paying attention that cramming ads into AI products is the next frontier in AI revenue streams.
First: You don't want to leak information about your users to advertising networks because it's going to leak, get back to your customers, they're going to figure out you're doing it and get really angry.
But second and more importantly - it's a much better business model to collect that data for yourself, keep it in house and then you control how you use that data to target ads which gives you a massive competitive advantage in selling ads because you have unique targeting data.
The way meta does this now is the model, they don't give the advertiser a list of the people you're going to show the advert to, the advertiser gives you a list of characteristics they want to hit and meta decides who those people are.
Apparently they have not learned, or they learned the wrong lesson. I do not think they leak it intentionally as hoarding is typically more profitable than selling but anyone who has been in the industry for some time knows that the move fast and break things attitude has caused enormous amounts of data leaks.
And then they will keep using your digital drugs like nothing happened and forget about the whole thing. So watch out, executive!
> But second and more importantly - it's a much better business model to collect that data for yourself, keep it in house and then you control how you use that data to target ads which gives you a massive competitive advantage in selling ads because you have unique targeting data.
Perhaps, but getting to the saturated market on this might be something LLM labs simply don’t have time for. They are haemorrhaging money, no path to profitability and OpenAI especially has made ridiculous promises on data centre -spending for the coming year. They need money now.
Not only this, many will continue to help the company by bullying anyone who decides to stop using the drugs.
The lesson they've learned is that if the profit from an activity is X, and the sanctions/reputational hit costs less than X, it's full steam forward.
It's not a 'leak'. It's in the T&Cs you agreed to. This is the business plan.
in this sense, any individualized incompetence is functionally malice-by-omission by organizations, especially ones as large and well-funded as LLM model providers
I would still bet they aren’t making any money from data leaking to advertisers
We have all become Milhouse now.
People love inequality when it's a celebrity they adore living large. Someone who hasn't scammed them and has demonstrably improved their life and the lives of others in a tangible way.
All we've chosen is free markets.
Unless you become a target of the government. Then people a lot worse than any teacher you ever had will be looking through it, and they, unfortunately, do care.
You are currently a cost to them¹ - using your data as free training/refinement, and potentially selling², it is a way to offset that a little.
----
[1] https://isaiprofitable.com/ - some of the green bars are creeping up a bit, but not much unless you count the shovel sellers
[2] sorry, leaking³
[3] Though it could of course be incompetence rather than malice, a mistake they are not actually making anything out of, as per Grey's amendment⁴ to Hanlon's razor.
[4] Any sufficiently advanced incompetence is indistinguishable from malice.
Access through: https://ai.ivx.run/chat/
I have recently noticed that e.g. ChatGPT, when used from a web browser, periodically sends unfinished prompts to their servers, namely to the `conversation/prepare` endpoint, without waiting for the user to actually send it.
This partial prompt data might potentially be used to "pre-warm" some kind of cache.
But it may also be used to track the user's writing cadence, error correction style and evolution of their stub ideas as they are being formulated into a prompt. I would assume that such data could also be sold to the advertisers.
OpenAI’s statements in response to the Millenium Prize (and related) disputes I think are a pretty obvious example of this in practice. One man’s “user prompts” is not another’s “reasoning trace scratchpad”.
This comment by Falserum on the mathematics research post articulates it well:
https://news.ycombinator.com/item?id=49649992
I'm pretty sure it is used for this; but rather than for anything nefarious, my guess is that this info is then fed to a classifier model to ensure that users of ChatGPT-the-service (as opposed to the OpenAI inference API) are actual humans, rather than agents trying to circumvent having to pay API pricing.
Sounds good.
What happened? Oh no!
How terrible! That’s just, that’s just awful!
How terrible! Oh no!