What You Share with AI Isn’t Always Private: A Business Owner’s Guide to Protecting Sensitive Data
Someone on your team is using AI right now.
Maybe it’s you. Maybe it’s your office manager, your bookkeeper, your top salesperson. They found a tool that makes their job easier, and they’re using it — probably without thinking twice about what happens to the information they type in.
That’s not a criticism. It’s just the reality of how AI adoption has spread through the American workplace. People discovered these tools were useful, and they started using them. Policies came later. Understanding came even later than that.
Imagine one of your employees is drafting a follow-up email to a client. They paste the client’s name, account balance, and a quick summary of the situation into ChatGPT to get a polished draft. They get a great result in ten seconds.
What they don’t realize is that the conversation may now be retained by OpenAI, potentially used to train future AI models, sitting on servers they have no access to or control over.
The client never consented. Your employee didn’t know. And there’s no way to take it back.
This article will tell you exactly what’s happening when someone types into an AI tool, what the real risks are, and what you can do about it — without giving up the productivity gains that made AI worth using in the first place.
Two Very Different Things
Before anything else, you need to understand a distinction that changes everything: inference versus training.
Most people don’t know it exists. And not knowing it leads to both overconfidence and unnecessary fear.
Inference is what happens in real time. You type a prompt. The AI processes it and generates a response. Your words are used to produce an answer, and that’s it — the model’s knowledge doesn’t change because of your conversation. It’s not getting smarter from your message in the moment.
Training is different. Training is how AI models learn in the first place — and how they keep improving over time. It involves feeding large amounts of data into the model to update its internal parameters.
The concern with user data is this: some AI companies use conversations from their consumer-facing products to improve future versions of their models. Your input from last Tuesday could influence how the model behaves for everyone next year.
Your data doesn’t end up sitting in a searchable database that other users can pull from. The risk isn’t that someone will type your client’s name into ChatGPT and see your exact conversation. The risk is more indirect — training shapes how the model responds in ways that reflect patterns from the data it learned from.
But that’s still a real risk, especially when the data involved is sensitive, identifiable, or proprietary.
The other thing worth knowing: opting out of training is forward-looking, not retroactive. If your data has already been used to train a model, flipping a privacy toggle today doesn’t remove it. The model doesn’t forget.
This is one of the most important and least-understood facts about AI data privacy, and we’ll come back to it.
What the Research Actually Found
In 2025, researchers at Stanford University’s Institute for Human-Centered AI published a study examining the privacy policies of six major U.S. AI companies: Amazon, Anthropic, Google, Meta, Microsoft, and OpenAI.
They analyzed 28 documents per company and applied the methodology of the California Consumer Privacy Act — the most comprehensive US privacy law, which all six companies are required to follow.
The findings were blunt.
All six companies use consumer chat data to train their models by default. Some retain that data indefinitely. Some allow human reviewers to read user conversations for training purposes. And some — but not all — claim to de-identify personal information before using it, though the study found that de-identification practices vary and are not consistently applied.
The lead researcher, Jennifer King, put it plainly: if you share sensitive information in a dialogue with ChatGPT, Gemini, or other major AI tools, it may be collected and used for training — even if you uploaded the information in a separate file rather than typing it directly.
That’s a significant statement. And it applies to everything your employees are putting into these tools every day.
A Lesson From Samsung
In March 2023, Samsung Semiconductor allowed its engineers to use ChatGPT. Within twenty days, three separate incidents of proprietary data exposure had occurred.
In the first, an engineer pasted source code from a semiconductor database into ChatGPT to fix errors.
In the second, an employee entered code used to identify defective equipment for optimization suggestions.
In the third, an employee recorded a confidential internal meeting and fed the transcript to ChatGPT to generate meeting notes.
Samsung had no enterprise agreement with OpenAI. No NDA. No data residency controls. No ability to request deletion from the model after the fact.
The engineers weren’t trying to cause harm. They were trying to work more efficiently . They just had no idea what happened to the data they typed in.
The fallout was immediate. Samsung banned ChatGPT and all similar AI tools on company devices. JPMorgan Chase and Amazon independently moved to restrict employee AI usage around the same time. Samsung later announced plans to build its own internal AI system to avoid routing proprietary data through third-party platforms entirely.
This scenario didn’t happen because Samsung had careless employees. It happened because there was no policy, no training, and no shared understanding of the risk.
That exact situation is playing out in businesses across the country right now — just without the media coverage.
The Data You Should Never Enter
Not everything carries the same risk. General business writing, publicly available research, marketing copy, anonymized examples — all of that is fair game for most AI tools.
The following categories are different. Keep them out of consumer-grade AI tools entirely.
Personally identifiable information (PII). Customer names combined with addresses, Social Security numbers, account numbers, dates of birth, or any combination of data that could identify a specific individual. Entering a customer’s PII into a consumer AI tool may violate your own privacy policy commitments to that customer.
Protected health information (PHI). Patient names, diagnoses, treatment plans, insurance details, or anything that could identify a patient. Using a consumer AI tool that doesn’t have a signed Business Associate Agreement (BAA) with your organization to process PHI is, in most cases, a HIPAA violation. Penalties range from $100 to $50,000 per violation.
Financial data. Bank account numbers, routing numbers, credit card details, customer transaction histories, or internal financial projections. This data is highly actionable for fraud. Depending on your industry, sharing it with unauthorized third-party systems may also create regulatory liability.
Trade secrets and proprietary business information. Unreleased product designs, source code, business strategies, pricing formulas, acquisition targets, supplier lists. Once this information leaves your network without proper contractual protection, you may have no legal recourse if it’s used against you.
Legally protected communications. Anything covered by attorney-client privilege, an NDA, or a confidential settlement. Entering privileged communications into a third-party AI tool may constitute a waiver of privilege — permanently.
Employee information. Compensation details, performance reviews, medical accommodations, background check results, Social Security numbers. Employee data carries its own set of legal obligations under state privacy laws and employment law.
The Plan Tier Matters More Than the Brand
Here’s the practical truth that most business owners don’t know: the brand of AI tool you use matters far less than which plan tier you’re on.
Consumer plans — free or low-cost personal accounts at ChatGPT, Claude, Gemini, and Microsoft Copilot — train on your conversations by default.
Business and enterprise plans don’t.
That’s the single most important distinction in this entire article. Let’s go through what it means by platform.
ChatGPT (OpenAI)
On free, Plus, and Pro consumer accounts, OpenAI uses your conversations to train future models by default. You can turn this off in Settings → Data Controls → “Improve the model for everyone.”
For one-off sensitive tasks, use Temporary Chat mode — it’s never used for training and is deleted within 30 days.
On ChatGPT Business and ChatGPT Enterprise plans, your data is not used for training by default. ChatGPT Enterprise is also eligible for a HIPAA Business Associate Agreement — but you have to explicitly request and execute one. The BAA is not automatic, and it is not available on the Business plan. Only Enterprise.
Claude (Anthropic)
This one catches people by surprise, because Claude built its early reputation on strong privacy defaults.
That changed in August 2025. Anthropic updated its consumer terms so that conversations on Free, Pro, and Max plans are now used for training by default — unless you opt out. Data can be retained for up to five years under the default settings.
To opt out: Settings → Privacy → turn off “Help improve Claude.” Doing so also drops retention to approximately 30 days.
For one-off tasks, Incognito Mode conversations are never used for training. The API and Claude for Work/Enterprise plans do not use data for training. A HIPAA BAA is available on the enterprise tier only, through a sales-assisted agreement.
Google Gemini
Consumer Gemini trains on conversations when “Gemini Apps Activity” is enabled — which it is by default. Turn it off in your Google account settings.
Even after opting out, conversations may be retained for up to 72 hours for processing before deletion. Google Workspace Business and Enterprise tiers don’t use your data for training.
Microsoft Copilot
Consumer Copilot started training on user data in October 2024 by default, with conversations retained for 18 months.
Microsoft 365 Copilot under a commercial or enterprise tenant is excluded from training by default. A HIPAA BAA is available under the Microsoft Online Services Data Protection Addendum for eligible commercial services.
The Critical Warning
“I use the paid version” is not a safe assumption.
A paid personal subscription — ChatGPT Plus, Claude Pro — is still a consumer plan. The data handling doesn’t change until you’re on a business or enterprise agreement.
Know which plan your employees are on.
How to Get Value Without Giving Away Data
Moving to a business plan solves a lot. But what about situations where you need AI help with something that touches real data, and you’re not ready to upgrade yet?
There are practical techniques that let you get useful AI output without exposing the actual sensitive information.
Anonymization and placeholders. Replace real names, numbers, and identifiers with generic labels before you type anything. Instead of “My client Jane Smith at Acme Corp owes $47,300,” try “My client [CLIENT A] at [COMPANY X] owes [AMOUNT].” The AI can still help you draft the email, structure the communication, or solve the underlying problem. It just doesn’t need the real details to do it.
Pseudonymization. When you need to maintain relationships between data points across a longer document or conversation, use consistent placeholder names that preserve the context without exposing real identities. “Dr. Sarah Smith, Oncologist” becomes “Senior Medical Specialist, Oncology Department.” “Chase Manhattan Bank” becomes “Major National Bank.” The structure stays intact; the identifiers go away.
Aggregation. Instead of entering individual records, summarize them. Rather than pasting a spreadsheet with 85 client names and balances, tell the AI: “I have 85 accounts receivable items over 90 days old, average balance of $2,400. Help me write a collections escalation policy.” Same help. No names.
Abstraction. Describe the structure of your problem, not the actual data inside it. “I have a spreadsheet with client name, contract value, and renewal date. I want to sort by renewal date and flag contracts under $10,000. How do I do that in Excel?” The AI doesn’t need your actual client list to answer that question.
Fictional examples. When you’re testing a workflow or building a process, create fictional examples that mirror the real pattern. Same structure, made-up details. You get to validate the approach without putting real data at risk.
One caution: don’t strip out so much context that the AI can’t help you meaningfully. Good anonymization preserves the structure and relationships the AI needs to be useful — it just removes the identifiers.
The Settings You Need to Change Now
These settings don’t fix everything, but they’re quick wins every AI user should take care of today.
Remember: opting out stops future training only. It does nothing about data already collected.
- ChatGPT: Settings → Data Controls → turn off “Improve the model for everyone”
- Claude: Settings → Privacy → turn off “Help improve Claude”
- Google Gemini: Google account settings → turn off “Gemini Apps Activity” (or “Keep Activity”)
- Microsoft Copilot (consumer): Check data controls in your Microsoft account settings
These settings are per-account. If your employees each have their own personal AI accounts, you have no visibility into whether those settings are on or off — and no way to enforce it centrally.
That’s a strong argument for migrating to company-managed business accounts.
What the Law Says
There’s no comprehensive federal AI privacy law in the United States right now. What exists instead is a patchwork — a mix of existing sector laws and rapidly expanding state regulations.
If you’re in healthcare, HIPAA is the most immediate concern. Consumer AI tools without a signed BAA cannot be used to process protected health information. New HHS-proposed regulations in January 2025 also introduced stricter cybersecurity requirements for HIPAA-covered entities.
The FTC has broad authority over unfair or deceptive practices, and it’s paying attention to AI. In September 2025, the FTC issued formal orders to seven major AI chatbot companies, demanding information about their data collection practices and how user data is used or shared.
The FTC has also clarified that non-HIPAA businesses collecting health information through apps or connected tools must comply with the Health Breach Notification Rule if there’s a breach.
At the state level, as of mid-2025, 20 U.S. states have enacted comprehensive data privacy laws. Eight of those became enforceable in 2025 alone. Several states are also developing AI-specific provisions — including requirements to disclose when personal data will be used to train AI models.
Colorado passed the first comprehensive state AI law in 2024, with key provisions taking effect in June 2026.
This varies significantly by state, business size, and industry. If you’re unsure of your obligations, talk to a qualified attorney before building AI into workflows that touch customer or employee data.
Your Business Needs a Policy
Individual awareness only goes so far.
According to a 2025 survey, 48% of employees have uploaded sensitive information to public generative AI tools, and 44% knowingly violated corporate AI policies. That second number is the one to pay attention to.
Nearly half of employees who broke the rules knew they were breaking them. They just didn’t have clear enough guidance to care, or the guardrails weren’t enforced.
A written AI acceptable use policy changes that. Not because it makes people suddenly compliant, but because it gives everyone a shared definition of what acceptable looks like — and it makes the consequences real.
A solid policy should cover:
- Approved tools — which AI products employees may use, and under what restrictions
- Data classification — a simple tier system defining what can go into any AI tool, what can only go into enterprise tools, and what can never go into any AI tool
- Prohibited use cases — explicitly name the categories: customer PII, PHI, financial account data, trade secrets, NDA-covered content
- Approved use cases — make it equally clear what you want employees to use AI for
- Account requirements — personal accounts or company-managed accounts only
- Incident reporting — what an employee should do if they realize they’ve shared something they shouldn’t have
- Consequences — what happens when the policy is violated
Build quarterly reviews into the policy from the start. AI provider terms change — Anthropic changed Claude’s consumer training default in August 2025 with relatively little fanfare. You won’t catch those changes unless you’re looking for them.
Frequently Asked Questions
Q: If I delete my chat history, does that remove my data from the AI’s training?
A: No. Deleting your conversation history removes your access to it, but data that has already been used for training cannot be retroactively removed. Opt-out settings prevent future training — they don’t reverse training that’s already occurred.
Q: Could another user see exactly what I typed into a chatbot?
A: Not directly. AI tools don’t work like a searchable database. Other users can’t retrieve your specific conversations. But data used for training can subtly influence the model’s future responses. The risk is indirect — patterns and information embedded in the model’s behavior, not your words showing up verbatim in someone else’s output.
Q: I pay for my AI subscription. Doesn’t that mean my data is protected?
A: Not automatically. Most paid consumer plans — ChatGPT Plus, Claude Pro — still train on your data by default unless you opt out. Training exclusions typically kick in at the business or enterprise tier, not the paid personal subscription level. Check your specific plan.
Q: We use AI through a third-party platform. Does the same risk apply?
A: It depends. Many third-party platforms access major AI providers through their APIs, and API usage is generally excluded from training by providers like OpenAI and Anthropic. However, the third-party platform has its own privacy policy that may allow additional data use. Review both.
Q: Our business is in healthcare. What do we need to know?
A: If you handle protected health information, you need a signed BAA with any AI provider before using their tool with that data. BAAs are only available on specific enterprise tiers — not on consumer plans, and not on most mid-tier business plans. If you’re in healthcare and using consumer AI tools with patient data, stop and consult a healthcare compliance attorney.
Q: We have NDAs with our clients. Does that protect us if we enter their data into an AI tool?
A: No. Your NDA governs your obligations to your client — it doesn’t bind the AI provider. Entering client data covered by an NDA into a consumer AI tool almost certainly violates the NDA itself, because you’re disclosing that data to an unauthorized third party.
Q: Is there any AI tool that doesn’t use my data at all?
A: Enterprise tiers of major providers don’t use your data for training by default. For complete data isolation, some organizations deploy AI models on their own private infrastructure or use locally hosted open-source models — meaning data never touches a third-party server. That approach is more technically complex and costly, but it eliminates third-party exposure entirely.
Q: How do I know if an employee has already shared sensitive data with an AI tool?
A: Without technical controls in place, you likely can’t. IBM’s 2025 Cost of a Data Breach report found that 86% of organizations are completely blind to their AI data flows. Enterprise AI tools with logging and audit capabilities are the primary technical solution. Clear policy and training reduce the risk in the meantime, but don’t guarantee visibility.
What Are You Going to Do About It?
Think back to that employee drafting the follow-up email. They weren’t trying to cause a problem. They were trying to do their job.
That’s how almost every AI data exposure happens — not through negligence or bad intent, but through a gap between what employees know and what they need to know.
The gap is closeable. It takes a policy, a clear set of rules about what data stays out of AI tools, the right plan tiers for your business, and a commitment to revisiting all of it as the tools and regulations continue to evolve.
The question isn’t whether AI belongs in your business. It almost certainly does.
The question is whether your team understands what they’re actually doing when they use it — and whether you’ve made it clear.
If you don’t know the answer to that, now is a good time to find out.
Sources:
- Security Boulevard: AI Acceptable Use Policy Guide
- Twig.so: Is Customer Data AI Training?
- Stanford News: AI Chatbot Privacy Risks Study
- The AI Career Lab: Does AI Train on Your Data?
- OpenAI: OpenAI Business Data Privacy
- Neowin: What AI Chatbots Collect
- BigID: Sensitive Data Types Explained
- Security Journey: Data Never Share with AI
- Justee.ai: What Not to Share with AI
- Kiteworks: IBM 2025 Data Breach Report
- The AI Career Lab: AI Business Associate Agreements
- Latitude.so: Privacy Risks in AI Prompts
- Caviard.ai: Anonymizing AI Prompts Guide
- Tanduo.io: AI Data Anonymization Guidelines
- ITECS Online: Create an AI Use Policy
- Stealthcloud.ai: Samsung ChatGPT Data Leak
- Sprypt: HIPAA AI Compliance 2025
- Federal Trade Commission: FTC AI Chatbot Inquiry 2025
- Federal Trade Commission: Health Breach Notification Rule
- American Bar Association: State Privacy Laws Expansion