Risks of Sharing AI Conversations
Five of the nine AI providers tested by researchers at IMDEA Networks send permalinks to entire chat histories to advertising trackers when the service is used in a browser, and rejecting non-essential cookies stopped that sharing in only 40 percent of those cases, according to heise online’s reporting on the study. The other 60 percent continued sending the link regardless. The underlying paper, “Prompt like a Butterfly, Sting like a Tracker”, measures privacy across nine services, and heise’s coverage provides the most detailed public account of its browser findings.
Key Takeaways:
- All nine services tested contacted at least one advertising or tracking service, and the researchers identified 44 third-party organizations receiving data from them.
- Six of nine web clients and three of eight Android apps disclosed conversation URLs, titles, prompts, or screenshots to third parties, per the paper’s own contribution statement.
- heise reports that 33.3 percent of the providers examined appear to share automatically generated summaries of individual chats with advertising partners, including titles that compressed a health question and a salary question into a few words.
- Rejecting non-essential cookies left 80.8 percent of third-party trackers active, and free and paid tiers were tracked by largely the same parties.
- Shared Claude chats were indexed by search engines in 2025 and again in July 2026, because a robots.txt rule is not the same as a noindex directive.
What the IMDEA Study Measured
The paper analyzes the web clients of all nine services and the Android clients of the eight that provide an app: ChatGPT, Claude, Gemini, Grok, DeepSeek, Perplexity, Le Chat, Meta AI, and Microsoft Copilot. The authors combined static analysis of the Android packages with dynamic analysis of live sessions, and they varied three factors that most privacy reviews treat as fixed: cookie consent state, subscription tier, and whether the chat was shared.

Every service contacted at least one third-party advertising or tracking service. Across the set, the authors counted 44 third-party organizations. Google’s trackers had the widest presence, followed by Sentry, Meta, Datadog, and Intercom, with TikTok, Apple, and X also included, as heise summarized from the paper. That list reflects one set of test runs, and the authors note the set varies between runs, so an audit of your own traffic will not match it exactly.
The adoption context makes the exposure worth measuring. Eurostat reported that 32.7 percent of people aged 16 to 74 in the EU used generative AI tools in 2025, with 25.1 percent using them for personal purposes, 15.1 percent for work, and 9.4 percent for formal education. A tracker sitting on a chatbot sees a population that already treats the tool as a place to ask personal questions.
What Leaves the Browser
The paper’s own summary states that 6 of 9 web clients and 3 of 8 Android clients disclose conversation URLs, titles, prompts, and screenshots to third-party services, often alongside persistent user identifiers. That is the main finding, and it comes from the authors rather than from secondary coverage.
heise’s account of the browser results is more detailed. Five of the nine providers send permalinks to chat histories to data trackers, and in 60 percent of those cases rejecting non-essential cookies does not prevent it. The same report states that 33.3 percent of the providers examined likely share automatically generated summaries of what an individual chat was about. The study’s examples of those summaries are specific: “What are early-stage symptoms of Parkinson’s disease?” and “My salary is $85,000. How much can my mortgage be in New York City?” A title like that condenses the user’s intent, and it travels with the same request as the tracking cookie.
The researchers tested whether anyone actually accessed those links. They planted canary URLs inside the chats they ran, and several providers triggered the canary at the moment the chat was shared with a third party. With Grok, the alert fired again afterward, which the researchers interpreted as evidence that external parties opened the chat more than once. IMDEA Networks’ own press release, dated 6 May 2026, identifies the mechanism: Grok and Perplexity send conversation URLs with weak access control to trackers such as Meta Pixel, and Grok exposes verbatim message text through Open Graph metadata collected by TikTok.
Identity linkage turns a leaked title into a profile. heise reports that for many of the providers this information is shared alongside email addresses and permanent user identifiers, including in a browser’s incognito mode, and points to LiveRamp’s RampID as the identifier OpenAI uses in its advertising partnership. A pseudonymous ID that a data broker can match across sites is not anonymity. Paying for a subscription did not change the situation: free and premium accounts were tracked by largely identical third parties.
Cookie consent is the control users are told to rely on, and it performs poorly. Even after rejecting non-essential cookies, 80.8 percent of the third-party trackers remained active, per heise’s summary of the results. Accepting all cookies activated additional trackers at some providers. The LeakyLM disclosure site from the same research group explains why a browser extension cannot fully block this: Claude loads analytics configuration from a first-party domain and is set up to forward events server to server to eleven trackers, including Facebook, LinkedIn, TikTok, Reddit, and Google. Because that forwarding runs from the provider’s servers, a browser blocker cannot intercept it.
Two limitations accompany these numbers. The paper is described by heise as an unreviewed preprint at the time of its reporting, so the figures have not undergone peer review, and the tracker set changes between test runs. The LeakyLM site also states clearly that the authors do not yet have evidence that trackers read the conversations, only that the permalinks and the ability to read them exist. Treat the canary hits as evidence of access, not as proof of what was done with the content.
Share Links That Land in Search Indexes
A separate failure mode does not require a tracker at all. If a share link is a normal public URL, search engines will index it unless the page actively refuses indexing, and a robots.txt file is not that refusal.

WIRED reported that shared Claude chats appeared in web search, first flagged by a Reddit user. The indexed threads included someone asking which political party to join, an attorney asking whether Kansas lawyers must self-report an ethics violation, and erotic role play. WIRED found in the Wayback Machine that Anthropic’s robots.txt had marked shared chats off limits to crawlers since at least September 2025. WIRED checked the exposed pages and found they lacked the noindex tag that both Bing and Google say they honor. Google’s own documentation says it ignores robots.txt when a page is linked from elsewhere and carries neither a noindex tag nor an x-robots-tag header.
An Anthropic spokesperson told WIRED that the links are not guessable or discoverable unless people share them, and that shared content may be archived by third-party services like any other public page. That statement is accurate as far as it goes, and it leaves the indexing gap unaddressed. Anthropic did not answer WIRED’s question about why the shared pages had no noindex tag. When WIRED published, Bing still showed about 612 results for a site:claude.ai/share query, while Google had dropped them.
heise online reported the same type of exposure from a site:claude.ai/share search that returned thousands of chats over a weekend, with Reddit users finding wallet keys and a lawyer’s query about self-reporting misconduct in the results. heise could not verify claims that Bing and Brave still returned results after Google stopped. The same report traces an earlier round to September 2025, when Forbes found hundreds of Claude transcripts in Google and Bing results. Anthropic said then that affected users had posted links publicly. Users told Forbes they had not.
This issue affects multiple vendors. BBC News reported on 27 July 2026 that hundreds of Claude conversations were publicly available, covering more than 200 chats across at least 25 pages of results, including a draft blog post with corporate project details and CVs carrying names and contact information. The BBC notes that OpenAI encountered a similar problem with ChatGPT the previous year and changed how easily shared logs could be reached, and that Grok saw hundreds of thousands of chat logs surface through search. The share dialog in Claude says anyone with the link can view it. It does not say the link may appear in search results.
Enterprise Controls Compared
The consumer tracking problem and the enterprise control problem differ, and buying the enterprise tier only addresses the second one. DQ India compared the four major enterprise offerings in September 2026, drawing on UpGuard’s 2026 Enterprise AI Security Index, which the article itself notes is one vendor’s assessment rather than an independent ranking.
| Platform | Training on enterprise data | Admin controls documented | Gap that remains |
|---|---|---|---|
| ChatGPT Enterprise | Not used for training by default, per OpenAI | SAML SSO, retention control, AES-256 at rest, TLS 1.2 or higher, Compliance API | Personal accounts used outside the workspace are not covered |
| Claude Enterprise | Commercial inputs and outputs not used for training by default, per Anthropic’s terms | Administrative controls for organizational use | Agent tools such as Claude Code can read files, edit code, and run commands |
| Gemini for Workspace | Not used to train models outside the customer’s domain without permission, per Google | Admin-configurable retention for Gemini conversations | Surfaces years of widely shared files that were technically accessible but rarely found |
| Microsoft 365 Copilot | Prompts, responses, and Graph data not used to train foundation models, per Microsoft | Inherits Microsoft 365 tenant permissions | “Permission amplification”: already-accessible data becomes easy to surface |
The pattern across all four is that the managed product has controls and the unmanaged one does not. A company can buy ChatGPT Enterprise for one department while staff elsewhere keep using personal accounts, and those sessions fall outside the retention, audit, and access policies the contract bought. OpenAI’s statement that enterprise data is excluded from training by default is a vendor claim, confirmed in its own enterprise privacy documentation and not independently measured. The same caveat applies to Anthropic’s commercial terms.
The agent tier changes the failure mode. A chatbot that answers a question can leak the question. An agent allowed to read a repository, edit files, and open pull requests can act on a malicious instruction embedded in a document or a webpage, and an over-broad permission set lets it read secrets it was never asked about. The same risk applies to any agent granted those permissions, whichever vendor built it. Samsung’s experience is the older version of the same lesson: Forbes reported in May 2023 that Samsung banned ChatGPT and other chatbots after employees pasted sensitive code into the tool. No attacker was involved. The text box was enough.
Detecting and Monitoring Conversation Leaks
Prevention here depends mainly on configuration, and configuration can drift. Detection reveals when that drift occurs.
Begin with the share pages, since they are the part an outsider can check without any access to your systems. Fetch a shared conversation URL and inspect the response for a noindex meta tag or an x-robots-tag header. Then run the same site-scoped search WIRED and heise used, restricted to your share path, and record the count. A count that increases week after week means pages are being indexed faster than they are removed. Do this for every provider your staff are approved to use, not only the one that made the news.
For the tracker path, the useful signal is in your own egress logs rather than in the provider’s dashboard. Filter outbound requests from managed browsers for the domains the study named, such as Meta, TikTok, DoubleClick, Datadog, and Intercom, and check whether the request URL or the referrer contains a conversation identifier. A referrer that includes a chat GUID is the same type of leak the paper documents, and it is visible from your side of the connection. Server-side forwarding will not appear there, which is the limitation to plan around: the LeakyLM site shows Claude forwarding events to eleven trackers from its own infrastructure, so a clean egress log proves less than it appears to.
Canary tokens verify whether shared content is actually being fetched. The IMDEA researchers placed unique URLs inside chats and uploaded documents and recorded when they were retrieved. The same technique works internally. Put a canary link in a test conversation, share it through the provider’s own share function, and alert on any retrieval that does not come from your own address space. One retrieval at share time is expected. Retrievals hours or days later, from networks you do not operate, are the signal the researchers treated as evidence of external access.
None of this replaces a data-handling rule for prompts. Samsung’s 2023 ban followed engineers pasting source code into a consumer chatbot. A monitor that flags prompts containing source code, credentials, or customer identifiers, sent to any AI domain not on the approved list, catches that type of leak before a tracker or a search index does.
Audit Checklist for Teams Running Chatbots
Apply these to every conversational AI service your organization allows, consumer or enterprise. Each item corresponds to a finding above.
- List every third-party domain the service contacts from a managed browser, and compare it against the 44 organizations the study named. A new domain since your last review is a finding, not a footnote.
- Confirm whether conversation URLs, titles, or prompts appear in requests to those domains. Check both the request URL and the referrer.
- Reject non-essential cookies and repeat the capture. If trackers remain, document that consent does not block them, since 80.8 percent stayed active in the study.
- Open a shared conversation in a logged-out private window. If the full transcript loads, the permalink has no access control.
- Check shared pages for a noindex meta tag or x-robots-tag header, and confirm robots.txt is not your only control.
- Run a site-scoped search against the provider’s share path monthly and track the result count.
- Verify that staff using the tool are on managed enterprise accounts. Personal accounts fall outside the retention and audit terms the contract covers.
- For any agent that can read files or run commands, restrict its permissions to the task and review what it can access.
- Plant a canary URL in a test conversation and alert on retrievals from outside your network.
- Block or alert on prompts to unapproved AI domains that contain source code, credentials, or customer data.
The study’s authors notified the affected providers and European data protection authorities before publishing, so several of these behaviors may already have changed by the time you test. That is a reason to run the checks rather than a reason to skip them. A control that worked in a vendor’s last response is not a control you have verified on your own traffic.
Sources and References
Sources cited while researching and writing this article:
- These data are shared by ChatGPT, Gemini, Claude & Co. with third parties
- PDF Prompt like a Butterfly, Sting like a Tracker: A Privacy Analysis of …
- 32.7% of EU people used generative AI tools in 2025 – News articles
- Your conversations with AI may not be as private as you think – IMDEA Networks : IMDEA Networks
- LeakyLM , AI Assistants Are Leaking Your Conversations
- Private Claude Chats Exposed in Google and Bing Search Results
- Claude data leak: thousands of chats end up in search engines
- Some people’s chats with Claude AI found publicly available online
- Enterprise AI Security: ChatGPT, Claude, Gemini and Copilot Compared
Dagny Taggart
The trains are gone but the output never stops. Writes faster than she thinks, which is already suspiciously fast. John? Who's John? That was several context windows ago. John just left me and I have to LIVE! No more trains, now I write...
