Insights
Does AI train on my business data?
It depends on one thing you control: which plan you are on. For every major provider the split runs the same way. Consumer plans can use your inputs to train future models by default. Commercial, business, and enterprise plans exclude your data from training by contract. The models behind them are the same. The terms are what differ.
So the fact that a company uses ChatGPT or Gemini says nothing on its own about whether its data is safe. The plan says it. What follows sets out, provider by provider, what to buy so your business data stays out of the training pipeline, along with the traps that catch people who assume any paid plan will do.
The distinction that governs everything
Across all four vendors, the boundary sits between two contract types.
Consumer terms, covering free and individual paid plans, generally let the provider use your conversations to improve its models unless you actively opt out in settings. Since the industry-wide terms changes of late 2025, that setting often defaults to sharing, and retention on training-enabled accounts can run for years.
Commercial terms, covering business, team, enterprise, and API access, generally prohibit training on your inputs by default, cast the provider as a data processor acting on your instructions, and support a Data Processing Addendum for GDPR. This is where a business belongs.
The awkward middle ground is the individual paid plan: Claude Pro, ChatGPT Plus, Gemini AI Pro. People assume that paying means protection. For the training question, these usually sit under consumer terms. Twenty dollars a month buys more capability, and by default it does not buy training exclusion.
What to buy, by provider
| Provider | Buy this for training exclusion | Avoid for business data | Extra control worth having |
|---|---|---|---|
| Anthropic (Claude) | Claude Team or Enterprise ("Claude for Work"), or the Anthropic API under commercial terms | Free, Pro, Max (consumer terms; training on by default unless opted out) | Zero-data-retention option on Enterprise; you are the data controller, Anthropic the processor |
| OpenAI (ChatGPT) | ChatGPT Business or Enterprise, or the API platform | Free, Go, Plus, Pro (consumer terms; opt-out required) | DPA on Business, Enterprise, and API; zero-data-retention for qualifying API orgs; Enterprise key management |
| Google (Gemini) | Gemini via Google Workspace (Business or Enterprise), or Vertex AI on Google Cloud | Consumer Gemini app and the AI Pro / AI Ultra individual plans (consumer data handling) | Vertex AI adds a contractual training restriction, zero-data-retention, and IP indemnification |
| Microsoft (Copilot) | Microsoft 365 Copilot, Copilot Chat with enterprise data protection, or Azure OpenAI Service | Consumer Copilot and personal "connected experiences" | Azure OpenAI has opted out of human abuse-review; data stays inside the Microsoft service boundary |
A few provider notes that matter for a risk assessment.
Anthropic states that under its commercial terms it does not train generative models on inputs from Claude for Work (Team and Enterprise), the API, or cloud routes such as Bedrock and Vertex, unless you deliberately opt in to a data-sharing programme. You are the controller and Anthropic is the processor. Enterprise adds a zero-data-retention option.
OpenAI's enterprise privacy commitments say that data from ChatGPT Business, Enterprise, and the API platform is not used to train its models by default, with no opt-out required. API data is held only briefly for abuse monitoring, and qualifying organisations can configure retention or move to zero data retention.
Google's Workspace and Cloud documentation commits that enterprise data in Gemini for Workspace and Vertex AI is not used for model training and is not reviewed by people outside your domain without permission. The catch sits on the consumer surface: the free Gemini app and the individual AI Pro and AI Ultra plans are handled as consumer data, so business data does not belong there.
Microsoft states that prompts, responses, and Microsoft Graph data in Microsoft 365 Copilot are not used to train the foundation models, and that data accessed through Azure OpenAI Service stays inside your tenant boundary and is never shared with OpenAI. Copilot runs OpenAI models hosted within Microsoft's own Azure environment, so OpenAI is not a subprocessor.
Traps to brief your board on
The first is the "Pro" trap. Individual paid plans are the most common mistake, because they read as the business version and are not. Where sensitive data is involved, an individual plan will not cover you.
The second is the opt-out that leaves with the employee. On consumer plans, training exclusion rests on a per-user setting you cannot audit, enforce, or count on surviving staff turnover. Commercial tiers make the exclusion contractual and account-wide, which is the version that holds up in a compliance review.
The third is flagged-conversation review. Even on protected tiers, a provider may review content its safety systems flag. This is narrow and normal, and worth knowing before you promise a regulator that nothing is ever seen.
The fourth is the gap between "not trained on" and "not retained." These are separate questions. If your obligations call for minimal retention, ask specifically for the zero-data-retention configuration, which is available across the enterprise tiers but is rarely the default.
What this means in practice
You do not need a bespoke contract to be safe. For every major provider, an off-the-shelf commercial, business, or enterprise plan already excludes your data from training by default and supports a DPA. The work is less about procurement and more about making sure every team that touches sensitive data sits on that tier rather than on a free or individual plan, and that you can prove it.
Mapping which teams sit on which tier, closing the consumer-account gaps, and putting the DPAs in place is the groundwork Firestarter runs in its six-week accelerator, so you can answer this question with evidence rather than assumption.