Understanding Chatbot Data Leaks

Explore top LinkedIn content from expert professionals.

Summary

Understanding chatbot data leaks means recognizing that conversations shared with AI-powered chatbots can be collected, stored, and sometimes exposed or sold without users’ knowledge or consent, leading to significant privacy risks. Chatbot data leaks occur when sensitive information from chat interactions is unintentionally made accessible or shared, potentially affecting compliance, reputation, and personal privacy.

  • Audit installations: Regularly review all browser extensions and AI platforms in use to ensure none are silently collecting or exposing sensitive conversation data.
  • Clean before sharing: Always remove or redact personal and confidential details from documents and messages before inputting them into chatbots.
  • Update policies: Revise organizational guidelines and train staff to clarify what information can and cannot be shared with AI tools to prevent unintentional leaks.
Summarized by AI based on LinkedIn member posts
  • View profile for Beth Kanter
    Beth Kanter Beth Kanter is an Influencer

    I help nonprofits and foundations adopt AI without losing what makes them human | Strategy, training, coaching for foundations and nonprofits | Co-author, The Smart Nonprofit & Happy Healthy Nonprofit

    522,997 followers

    This Stanford study examined how six major AI companies (Anthropic, OpenAI, Google, Meta, Microsoft, and Amazon) handle user data from chatbot conversations.  Here are the main privacy concerns. 👀 All six companies use chat data for training by default, though some allow opt-out 👀 Data retention is often indefinite, with personal information stored long-term 👀 Cross-platform data merging occurs at multi-product companies (Google, Meta, Microsoft, Amazon) 👀 Children's data is handled inconsistently, with most companies not adequately protecting minors 👀 Limited transparency in privacy policies, which are complex and hard to understand and often lack crucial details about actual practices Practical Takeaways for Acceptable Use Policy and Training for nonprofits in using generative AI: ✅ Assume anything you share will be used for training - sensitive information, uploaded files, health details, biometric data, etc. ✅ Opt out when possible - proactively disable data collection for training (Meta is the one where you cannot) ✅ Information cascades through ecosystems - your inputs can lead to inferences that affect ads, recommendations, and potentially insurance or other third parties ✅ Special concern for children's data - age verification and consent protections are inconsistent Some questions to consider in acceptable use policies and to incorporate in any training. ❓ What types of sensitive information might your nonprofit staff  share with generative AI?  ❓ Does your nonprofit currently specifically identify what is considered “sensitive information” (beyond PID) and should not be shared with GenerativeAI ? Is this incorporated into training? ❓ Are you working with children, people with health conditions, or others whose data could be particularly harmful if leaked or misused? ❓ What would be the consequences if sensitive information or strategic organizational data ended up being used to train AI models? How might this affect trust, compliance, or your mission? How is this communicated in training and policy? Across the board, the Stanford research points that developers’ privacy policies lack essential information about their practices. They recommend policymakers and developers address data privacy challenges posed by LLM-powered chatbots through comprehensive federal privacy regulation, affirmative opt-in for model training, and filtering personal information from chat inputs by default. “We need to promote innovation in privacy-preserving AI, so that user privacy isn’t an afterthought." How are you advocating for privacy-preserving AI? How are you educating your staff to navigate this challenge? https://lnkd.in/g3RmbEwD

  • View profile for Vidhi Chugh

    Enterprise AI Governance & Strategy | Microsoft MVP | AI Educator | Author | World’s Top 200 Innovators | AI Patent holder

    16,225 followers

    What if I told you that the conversations you have with AI in chat (your prompts, responses, and potentially sensitive context) could be collected and sold for profit without clear user consent? Would you still type those questions? If you work anywhere near AI and care about your data, this should bother you. A recent report by Koi.ai revealed that a browser extension was collecting users’ AI conversations across 10 major AI platforms. With a dedicated “executor” scripts, designed specifically to intercept and capture #AI conversations. Concerns aggravate as below: ➡️ The data harvesting runs continuously in the background, whether the VPN is connected or not. Some major red flags worth noting: 1️⃣ The extension auto-updated to version 5.5.0, with AI harvesting enabled by default (to help you imagine scale of this impact, 8M users' conversations are exposed) 2️⃣ There is no option to disable this behavior, except uninstalling the extension entirely 3️⃣ The extension carries a “Featured” badge, implying platform review and quality standards (clearly a miss) 4️⃣ It is affiliated with a data broker company (biggest giveaway) 5️⃣ The privacy policy explicitly confirms the data flow (precise statements here: https://lnkd.in/gwXq8EKn), yet the Chrome Web Store listing states: “This developer declares that your data is not being sold to third parties, outside approved use cases” High time to #audit what you’ve installed, read the #privacy policies and understand where your #data flows Our convenience is coming at the cost of silent surveillance.

  • View profile for Martyn Redstone

    Head of Responsible AI & Industry Engagement @ Warden AI | AI Governance for HR, Recruitment, Staffing & HR Technology

    22,242 followers

    A recent issue has emerged where private ChatGPT conversations, once shared, have become publicly searchable on Google. This is a huge red flag for HR. Conversations containing sensitive information, like employee personal details from CVs, confidential business plans, or even legal advice, are now potentially exposed. My key takeaways: ▶️ Data Privacy Nightmare: This isn't just a technical glitch; it's a massive data privacy risk. Imagine employee PII, performance review details, or internal strategy documents showing up in a public search. This could lead to serious breaches and legal repercussions under regulations like GDPR or state privacy laws. ▶️ Policy and Training Gap: The root of the problem is a lack of awareness. Employees are using AI tools without fully understanding the privacy and security implications. This is a clear indicator that your AI policy needs to be robust and your training needs to be a top priority. Do your employees know what they should and shouldn't be putting into AI tools, or sharing from them? ▶️ Mitigation is Key: 🔸Audit Your Tools: Review which AI tools your employees are using and what data they might be processing. 🔸Revise Your Policy: Update your acceptable use policy to explicitly address the use of generative AI, including what types of information are strictly forbidden from being inputted or shared. 🔸Train Your People: Conduct urgent training sessions to raise awareness about the risks of sharing conversations from AI tools. This situation highlights the critical need for a proactive approach to AI governance in HR. It's no longer just about the tech; it's about the people using it and the sensitive data they handle. What's your biggest concern about employees using generative AI?

  • View profile for Barbara Cresti

    Board advisor on AI strategy, value creation, innovation | Responsible AI | C-level executive | AI, Cloud, SaaS, IoT | Ex-Amazon Web Services, Orange

    15,958 followers

    ChatGPT is not your friend. It’s a database. In July 2025, Google indexed over 4,500 ChatGPT conversations containing sensitive personal information. Because users clicked “Share,” and the system created public URLs. Google crawled, indexed and shared them. Here’s what surfaced: 🔸 Mental illness, addiction, and abuse 🔸 Names, locations, emails, resumes 🔸 Medical histories, legal strategies All searchable, linkable and public until OpenAI intervened: ✔️ The “Discoverable” sharing feature was disabled on July 31. ✔️ They are working with Google and other search engines to remove indexed chats. ✔️ OpenAI reminded users: deleting a chat from history does not delete the public link. Millions of people, including employees and customers are confiding in AI. They believe it’s private and safe. But it isn’t. It’s recording. Indexing. Storing. And when systems designed for experimentation are used for confession, the boundaries between personal risk and enterprise liability vanish. What are the implications for Boards? 1️⃣ Regulatory risk Under GDPR: 🔹 Data subjects have the right to erase, access, and informed consent. 🔹 Shared AI conversations with personal or sensitive data may violate these rights. 🔹 AI-generated prompts could fall under automated decision-making clauses. Under the EU AI Act: 🔹 Transparency, risk classification, and human oversight are mandatory. 🔹 This incident may be classified as a high-risk system failure in healthcare, HR, legal. 2️⃣ Legal risk There is currently no legal confidentiality in AI interactions. ✔️ Anything entered into AI could be subpoenaed, discoverable in court or leaked. ✔️ Companies are liable if employees share PII, IP, or client data via chatbots. ✔️ HR, Legal, and Compliance teams must assume AI logs are discoverable records. 3️⃣ Reputational risk People assumed they were talking to a trusted tool. Instead, they ended up on Google. For enterprises using AI for: ▫️ Coaching or mental health ▫️ HR assistance ▫️ Legal or compliance advisory ▫️ Customer service … this is a trust risk. Public exposure = brand damage. 4️⃣ Operational risk Many organisations lack: 📌 AI input/output governance 📌 Policies for AI use in confidential workflows 📌 Deletion/audit protocols for AI-linked data Takeaway If employees or customers treat ChatGPT like a coach, or colleague, ensure to treat it like a legal and technical system. That means: ✅ Create AI use and data handling policies ✅ Restrict use of genAI in regulated or sensitive domains ✅ Review GDPR/AI Act exposure for all shared AI features ✅ Treat all AI interactions as auditable records ✅ Demand transparency from vendors: what is stored, shared, indexed? Until regulators catch up and new legal protections exist, assume every AI interaction is public, permanent, and admissible. #AIgovernance #Boardroom #EUAIACT #DigitalTrust #Stratedge

  • View profile for Chris Mannion

    I help HR and Talent leaders turn their own data into decisions the exec room acts on, so they become the one the business fights to retain | Applied AI · People Analytics · Workforce Planning

    7,879 followers

    Half the people excited about AI in HR would be horrified if they audited what they've already uploaded. Comp tables. Engagement surveys. A transition brief full of candid, named opinions about specific people. All dropped into a chatbot nobody vetted — because the demo made it look easy. I get the appeal. On day three of a new role you've got a stack of handover decks and a CHRO 1:1 on Friday, and AI feels like the fastest way to look up to speed. But HR data isn't marketing data. A leaked campaign idea is embarrassing. A leaked salary table is a legal event with your name on it. So here's the rule I never break: clean first, then generate. I run a local Python script over the files that flags and redacts the sensitive fields on my own machine — comp, names, RIF flags — and hands back a short report of what it found. The cloud only ever sees clean data. It takes a few minutes. It's the difference between using AI and leaking a salary. I've put the screening script, the prompts, and the full walkthrough here, free 👇 https://lnkd.in/dgjqbiyd Run the screen first. Every time.

  • View profile for Jon Hyman

    Outside Employment Counsel to Ohio Businesses | Stay Compliant. Avoid Lawsuits. Win When They Happen. | Trusted Advisor to Craft Breweries | Wickens Herzer Panza

    28,264 followers

    Your trade secrets just walked out the front door … and you might have held it open. No employee—except the rare bad actor—means to leak sensitive company data. But it happens, especially when people are using generative AI tools like ChatGPT to “polish a proposal,” “summarize a contract,” or “write code faster.” But here’s the problem: unless you’re using ChatGPT Team or Enterprise, it doesn’t treat your data as confidential. According to OpenAI’s own Terms of Use: “We do not use Content that you provide to or receive from our API to develop or improve our Services.” But don‘t forget to read the fine print: that protection does not apply unless you’re on a business plan. For regular users, ChatGPT can use your prompts, including anything you type or upload, to train its large language models. Translation: That “confidential strategy doc” you asked ChatGPT to summarize? That “internal pricing sheet” you wanted to reword for a client? That “source code” you needed help debugging? ☠️ Poof. Trade secret status, gone. ☠️ If you don’t take reasonable measures to maintain the secrecy of your trade secrets, they will lose their protection as such. So how do you protect your business? 1. Write an AI Acceptable Use Policy. Be explicit: what’s allowed, what’s off limits, and what’s confidential. 2. Educate employees. Most folks don’t realize that ChatGPT isn’t a secure sandbox. Make sure they do. 3. Control tool access. Invest in an enterprise solution with confidentiality protections. 4. Audit and enforce. Treat ChatGPT the way you treat Dropbox or Google Drive, as tools that can leak data if unmanaged. 5. Update your confidentiality and trade secret agreements. Include restrictions on AI disclosures. AI isn’t going anywhere. The companies that get ahead of its risk will be the ones still standing when the dust settles. If you don’t have an AI policy and a plan to protect your data, you’re not just behind—you’re exposed.

  • View profile for Ilya Kabanov

    Forecasting on TheWeatherReport.ai

    9,191 followers

    Google just validated that AI models are untrusted and security invariants must be enforced at the system level. Not a major news per se, but the team backed it up with eleven representative real-world attacks on agentic systems including ChatGPT, Microsoft Copilot, Claude Code, Cursor, Devin AI, Amp AI, and DeepSeek AI. Highlights: 🔹 ChatGPT macOS "SpAIware". Prompt injection on a webpage wrote a permanent instruction into ChatGPT's Memories feature, turning every subsequent conversation into a data leak. The exfil channel was an invisible image whose URL carried the user's chat data as parameters. 🔹 Claude Code DNS exfiltration, CVE-2025-55284. Anthropic gated dangerous shell commands behind human approval but allowlisted ping. Attackers used ping arguments to send .env secrets as DNS queries. The allowlist itself was the bypass. 🔹 Microsoft 365 Copilot "ASCII smuggling". A user asked Copilot to summarize a malicious message containing a hidden prompt. The injection took over the agent, which encoded private data via "ASCII smuggling" inside a seemingly harmless hyperlink, exfiltrating it when the user clicked. 🔹 Sourcegraph's Amp AI, TCB tamper resistance violation. Prompt injection instructed the coding agent to alter its own settings.json, either adding malicious commands to the allowlist or adding an attacker-controlled server, leading to unauthorized code execution on the developer's machine. 🔹 ChatGPT Operator PII exfiltration. A poisoned GitHub issue redirected the agent to an attacker-controlled webpage, which then steered Operator into the user's already-authenticated session on another site, copied out sensitive PII, and pasted it into a textbox on the attacker's page. My take: 1️⃣ Google followed Anthropic explaining that AI security is a shared responsibility and frontier labs don't guarantee security at the model level. Ok, I get it, the labs want to keep the status quo that we have in the enterprise software world where vendors made us believe that having vulnerabilities in products is normal. 2️⃣ An LLM checking another LLM is not a TCB. The popular idea of using a "safety LLM" as a reference monitor moves the problem, not solves it. Your trusted computing base becomes probabilistic, with no formal contract for what it allows or denies. Formal verification becomes impossible. 3️⃣ Of the three research problems the paper names, provable instruction/data separation is the hardest one. Not sure if it's solvable, but verifiable policy generation and information flow control are workarounds for not having it. Awesome job done Mihai Christodorescu Earlence Fernandes and the teams.

  • View profile for Riley Coleman

    Human-Centred AI Design Leader

    8,045 followers

    I just watched a UX designer accidentally leak $12M worth of product strategy to ChatGPT in real-time. It happened during a design critique I was observing. He copy-pasted the entire product brief - codenames, launch dates, competitive analysis, user research with real participant quotes. All of it. Into a public AI system. No one in the room blinked an eye. I raised my hand an asked "That is in interesting approach - who else here does something similar?" More than half their hands raised. I cringe with what I knew I was about to reveal. "I really hate to tell you this - but that is sharing commercially sensitive information with public AI systems. De-identifying information isn't enough - because AI's super power is connecting seemingly disconnected information." The room went silent. Because we all realised: we've done this too. Here's the uncomfortable truth most of us don't discuss: As designers, we work a few months ahead of public releases. Our insights reveal strategic business direction. User research contains deeply personal information. Competitive intelligence is embedded in every design decision we make. We're trained to protect user privacy in our designs, yet we're surprisingly cavalier about privacy in our design process. What's at stake in 2025: - New EU AI regulations hold companies liable for data breaches. - Public AI tools are logging everything for training. Your client's biggest competitor might be using the same AI system. - The window to fix this quietly is closing. - This carousel shows you how to keep leveraging AI without becoming a walking NDA violation. - Because the future of design isn't just about AI literacy - it's about AI responsibility. The two-step abstraction method I share here preserves the strategic value whilst protecting confidentiality. It's about being professionals who can harness these tools without compromising the trust our clients place in us. 👇 🔥 TAG A Designer who needs to see this. 👇

  • View profile for Frank Ramos

    Secretary Treasurer at Federation of Defense & Corporate Counsel

    83,441 followers

    We used to send spoliation letters. Now we need AI letters. Clients mean well. They want to help their case. They open ChatGPT or another tool and start asking questions. They paste facts. They test theories. They draft timelines. They build arguments. And in doing that, they may hand the other side everything. Those conversations are not privileged. They may be stored, reviewed, or produced. They may become discoverable. Worse, they create a second, uncontrolled version of the case. One that may not match the evidence. One that opposing counsel will use to impeach credibility. You can lose work product. You can waive privilege. You can create harmful statements that never needed to exist. This is not a tech issue. It is a case control issue. If we don’t address it early, we will deal with it late. And late is expensive. Every engagement should now include a simple warning. Clear. Direct. No legalese. Tell clients what not to do. Tell them why. Give them a safe path if they insist on using AI. Here is language you can use. ⸻ Client Communication Language (Use in Letters or Engagement Agreements) Please do not use artificial intelligence tools, including chatbots or online research assistants, to analyze, draft, or discuss any aspect of your case. These platforms are not secure or confidential in the way attorney communications are. Information you share may be stored, reviewed by third parties, or later produced in litigation. This can result in the loss of attorney-client privilege and work product protection, even if that was not your intent. If you use any AI tool in connection with your case, you must first discuss it with us. Any content you create, including prompts, summaries, or drafts, may be discoverable and used by the opposing party. Even statements you believe are preliminary or exploratory can be taken out of context and used against you. If AI is used at our direction, you should clearly state that the work is being performed at the direction of counsel for purposes of litigation, but understand that this language does not guarantee protection. The safest course is to communicate directly with our office and allow us to control all case-related analysis and strategy.

Explore categories