Skip to main content

Which AI Providers Train on Your Data? A 2026 Comparison

Does your AI provider train on what you paste in? We read 21 policies. What each one says, how long they keep it, and what changes with your plan.

Sujith S
CTO, Product Owner, Zackriya Solutions
12 min readPrivacy & SecurityEnglish
Which AI providers train on your data? The three things this comparison covers for each provider: training, retention and human review, from 21 AI privacy policies reviewed

TL;DR

  • Most paid AI APIs say they don't use your data to train their models. Of the four meeting note-takers we checked, two say the same about your recordings, and two train on de-identified versions of them.
  • What happens after they receive the data is less straightforward. Retention, human review, regional privacy policies, and plan-specific terms can all change the picture.
  • The same provider can have different rules for its consumer app and API. Regional policies can matter too. OpenAI, for example, has separate policies for Europe, the US, and the rest of the world.
  • Retention periods are different across providers. Anthropic's consumer default is 30 days, but enabling model improvement can extend retention to up to five years.
  • Your data often goes further than the company you picked. AI services pass it to other companies for hosting, transcription or the model itself, and cloud note-takers often use several. You agree to the first company's terms, not theirs.
  • We reviewed the companies' own documentation between 16 and 24 September 2026. These are their published statements and commitments, not independent audit results.

Meeting transcripts contain more sensitive information than most teams realize.

A discussion about layoffs. A customer's complaint. A pricing figure someone didn't want shared. A hiring decision. An internal disagreement.

Once a meeting has been transcribed, all of that becomes searchable text. And sooner or later, someone is likely to paste part of it into an AI tool to get a summary.

That's where a simple question comes up:

What happens to that text after you send it?

Does the provider train its models on it? How long does it keep it? Can people review it? Where is it stored? And does the answer change depending on whether you're using a free app, a paid API, or an enterprise plan?

Those questions are more useful than simply asking, "Does this AI train on my data?"

Most AI providers say they don't use data sent through paid APIs to train their models. Meeting note-takers are more mixed: of the four we checked, two say they don't train on your recordings, and two train on de-identified versions of them.

That still doesn't tell you everything.

You need to know what happens to the data after it's collected, how long it's kept, which privacy policy applies to your account, and what protections are included in the plan you're using.

We went through 21 privacy policies and related documents to build the tables below.

A note before you read

Everything here comes from the companies' own documentation, reviewed between 16 and 24 September 2026.

These are statements made by the companies themselves. They are not audit results.

We haven't inspected their infrastructure or independently checked what happens inside their systems. Reading a privacy policy can tell you what a company says it does. It can't prove what happens behind the scenes.

So this isn't a ranking of which provider you should trust.

It's simply a summary of what each company currently says about how it handles your data.

What the providers say

The provider names link to our longer write-ups, or to the provider's own documentation where we don't have one.

The first five are providers where the rules are different for the consumer product and the API. It's an easy detail to miss when reviewing a provider.

ProviderDoes the consumer app train on your data?Does the paid API train on your data?How long your prompts and replies are kept
OpenAIYes, by defaultNoAPI: 30 days
Google GeminiYes, with human reviewNo on paid, yes on the free tierConsumer: 18 months
MistralYes on the free tierNoAPI: 30 days
PerplexityYes, by default on Free, Pro and MaxNoAPI: nothing kept
xAI GrokYes, unless you turn it offNoAPI: 30 days
Anthropic ClaudeNo, unless you opt inNoConsumer: 30 days, or up to 5 years
Azure OpenAINo consumer productNoAPI: only flagged samples, kept for human review. No period stated
AWS Bedrock(opens in new tab)No consumer productNoAPI: depends on the model. Zero retention available
GroqNo consumer productNo, by contractAPI: nothing, by default
Fireworks AINo consumer productNoAPI: nothing, with one exception
OpenRouterNo consumer productNoAPI: nothing unless you opt in
Together AINo consumer productNoAPI: stored by default
CohereNo consumer productYes, unless you opt outAPI: 30 days
DeepSeekYes, with a right to opt outNo exclusion statedNo limit published, either tier

About that last column. Each provider publishes retention for the tier that matters most to its own users, so we've labelled which one each figure refers to. Don't read an "API: 30 days" against a "Consumer: 18 months" as if they measure the same thing. Where a provider publishes both, the notes below give you the other one.

Eleven of those rows need a little more detail.

OpenAI. The 30 days covers abuse monitoring. Some features hold your data until you delete it yourself, and zero data retention doesn't cover them. The list includes Conversations, Assistants, Threads, Vector Stores, Files, fine-tuning jobs, evals and batches.

Google. Consumer activity deletes automatically at 18 months, and you can change that to 3 or 36 months, or switch it off. Chats picked for human review are kept up to three years, and Google says these "are not deleted when you delete your activity." The free API tier gets human review too.

Anthropic. 30 days is the stated default. If you allow Anthropic to use your chats to improve Claude, it "may retain your data in a de-identified format for up to 5 years in our model training pipelines." Incognito chats are excluded either way.

Perplexity. The API keeps nothing and isn't used for training. The consumer app is the opposite way round. AI data retention is on by default for Free, Pro and Max, and you switch it off under Account, then Preferences, then AI data retention. Some clients label the same control AI Data Usage. Turning it off only applies from that point on, not to what you sent before. Enterprise data is never used for training.

Mistral. Inputs and outputs are kept for 30 rolling days to monitor abuse, unless zero data retention is switched on. Two exceptions: the Agents API keeps data until you close your account, and fine-tuning data stays until you delete it.

xAI. The API position is clear: xAI "never trains on your API inputs or outputs without your explicit permission," with a 30-day encrypted audit window.

The consumer side works the other way round. Grok is an opt-out product: the control exists, and you have to go and switch it off. xAI's own FAQ doesn't say which way the switch starts, but X's help page describes Grok on X as something you opt out of. So treat it as on, and check it after you install rather than assuming.

The bigger catch is that there isn't one setting, there are three. On the Grok mobile app it's Settings, then Data Controls, then "Improve the model." On grok.com it's Settings, then Data, then "Improve the Model." Grok inside X is a separate control again, in X's own settings under Privacy and safety, then Data sharing and personalization, then Grok & Third-party Collaborators. Nothing we read says the three sync, so switch off each one you use. Private Chat is the exception: xAI says those conversations aren't used for training and are deleted within 30 days.

Groq. The commitment sits in the Services Agreement, not on a policy page. That makes it a contract term rather than a stated practice. More on why that matters below.

Fireworks. Nothing is stored for open models without your opt-in. The Response API is the exception: it defaults to store=True and keeps data for 30 days. Note that Fireworks states it doesn't log or store your prompts, rather than making a separate statement about training. Data it never keeps can't be trained on, but that is our inference and not their sentence.

Together AI. They don't train on your data by default. They do store it by default. Both are true. An admin has to switch storage off.

Cohere. The page doesn't use the word "default." It says you can "opt out from your prompts and generations being used to train Cohere models in your dashboard settings at any time," and that you set the toggle to "Off" to opt out. Worth knowing: the Enterprise Data Commitments apply to paying customers with a card on file. Trial API keys fall under the standard Terms of Use and Privacy Policy instead.

DeepSeek. Its privacy policy lists training its models as one use of your data, and it doesn't carve out the API. Its training explainer says only "a small portion" of training data may come from user input, de-identified first, and that you can opt out. Data is collected, processed and stored in the People's Republic of China. No deletion timeline is published, only "as long as necessary."

Then there's the local option, which isn't a policy at all. Ollama, LM Studio and llama.cpp run the model on your own machine. Nothing gets sent anywhere. There's no document to read and no commitment anyone has to keep.

Two things to take from that table

Training and retention are separate questions. A "no" on training tells you nothing about how long your text sits on someone's server. Together AI shows the gap clearly.

Human review is also more common than most buyers expect. Google says so for the consumer app and for the free API tier. Microsoft says so for flagged content. Several others say so for abuse enforcement. None of this is improper. It's just rarely what people have in mind when they read "we don't train on your data."

Four things the table can't tell you

The policy you read might not be the one that covers you

OpenAI publishes at least three privacy policies. Your region decides which one applies to you.

There's a Europe policy(opens in new tab) for the EEA, the UK and Switzerland. There's a US policy(opens in new tab). And there's a rest-of-world policy(opens in new tab) that points you to whichever of those two applies to you.

The company responsible changes too. It's OpenAI Ireland Limited if you're in the EEA or Switzerland. It's OpenAI OpCo, LLC in San Francisco if you're anywhere else, and that includes the UK. So someone in London reads the Europe policy, but the company responsible for their data is the US one.

Google does its own version of this, and says so plainly. The free Gemini API tier is used to improve Google's products, and may be read by human reviewers. That changes if you're in the EEA, Switzerland or the UK. There, Google applies the paid terms to everything, "even though they are offered free of charge."

So two colleagues on the same plan, using the same model, can have different rights. One sits in Dublin and one sits in Bangalore. We have engineers in both.

If your company works across regions, "we checked the privacy policy" isn't a finished answer. Which policy, and for which staff?

Two entrances in the same wall. The app your team opens trains by default. The API you approved does not.
At OpenAI, Google, Mistral, Perplexity and xAI, the consumer product and the paid API have different training defaults. Verified 24 September 2026.

Better privacy often costs more

Take one company and one model, then move up the price list.

ChatGPT Enterprise(opens in new tab) says business data isn't used to train OpenAI's models by default. It also adds admin controls, plus compliance and audit visibility. Consumer ChatGPT on a personal account has "Improve the model for everyone" switched on unless someone turned it off.

The same pattern shows up almost everywhere.

CompanyFree or consumerPaid or enterprise
OpenAIPersonal ChatGPT: trainsAPI, Business, Enterprise: doesn't, by default
GoogleGemini app, free API: trainsPaid API: doesn't
MistralFree tier: trainsPaid: can opt out. Enterprise: already out
PerplexityFree, Pro and Max: on by defaultEnterprise: never used for training
GranolaDe-identified data, opt out in settingsEnterprise: off by default

Read that column break and something becomes clear. Part of what an enterprise tier sells is the privacy promise itself.

We're not going to call that a scandal. Abuse monitoring, regional data stores and compliance APIs cost real money to run. Somebody has to pay for them. We run infrastructure too.

But it has a practical effect. The protections described in a vendor's enterprise documentation are often the ones you didn't buy. And the free tier is the one your team reaches for when they're busy.

A promise in a contract is stronger than a setting in a dashboard

These two often get treated as the same thing. They aren't.

A contract term binds the company. Groq's Services Agreement says Groq "is not permitted to use Inputs or Outputs for training or fine-tuning" without your instruction. Azure and AWS put their commitments in their terms as well. A term like that survives staff changes, and breaking it has consequences.

A dashboard setting is different. Cohere's Data Controls, Mistral's data-sharing setting and the consumer "improve the model" switches all work. They can also be changed by anyone who can log in. That includes someone new who doesn't know why the setting was there.

Both are real. They're just not equally strong. If the answer to "can we opt out?" is a screenshot of a settings page, that's weaker than a clause with a date on it.

Your data goes further than the company you picked

Most AI services don't do all the work themselves. They pass your data to other companies for hosting, for transcription, or for the model itself. Each of those companies has its own terms, and you usually never see them.

OpenRouter is the plainest case on the provider side, because routing to other companies is the product. OpenRouter doesn't store your prompts unless you opt in. But it says "each AI provider on OpenRouter has its own data handling policies for logging and retention." To its credit, it gives you a setting to avoid routing to providers that may train on your data.

Cloud meeting note-takers add more links to the chain. Granola's security page(opens in new tab) says it uses transcription providers "like Deepgram and Assembly" and AI providers "like OpenAI and Anthropic." So your audio goes to one company, your transcript goes to another, and your notes are stored on AWS.

Those suppliers have their own rules. Deepgram's terms(opens in new tab) say it may use customer content for "training and testing our Models," and that API customers can opt out "on a per-request basis." Granola says it doesn't allow third parties to use your data to train their models. That's a real protection, and it's Granola's to keep. It depends on an opt-out or a contract term being applied every time, in systems you can't see.

None of this means a company is careless. It's how most cloud software is built. But every extra company that handles your meeting is one more place it can leak from, and one more set of terms you're relying on without having read.

What about the meeting tools themselves?

The answers here are more mixed than on the provider side, and two of them are firmer than you might expect, which is worth saying clearly.

ToolTrains on your meetings?What it collects to workWhere it sits
FirefliesNo, and it says its vendors can't eitherParticipants, emails, meeting title, audio and video, meeting IDs, Voice DataUnited States and other countries
tl;dvNo, for general-purpose AI modelsRecordings, transcripts, participant emails including non-customers, IP, locationMainly EU. AI features may run in the US or EU
GranolaYes, on de-identified data. You can opt out in settings, and Enterprise is off by defaultRecordings, transcripts, calendar invites and body text, contacts, employment dataAWS, United States
OtterYes, on de-identified recordings and transcripts. No opt-out describedRecordings, OtterPilot screenshots, speaker IDs, calendar, contacts, locationAWS, United States
MeetilyNo. We never receive the recordingNothing from your meetingYour device

Three details from those policies are worth knowing.

Fireflies notes that its Voice Data "may be considered 'biometric identifiers' or 'biometric information' in some jurisdictions." It adds that its providers don't use this data to identify anyone, and that Fireflies never receives it on its own servers.

tl;dv names its sub-processors openly. That list includes AssemblyAI, ElevenLabs, Anthropic and Google Vertex, with hosting in Germany and Finland. That's more disclosure than most tools in this category offer. Its wording on training is specific: customer content "is not used by tldx Solutions GmbH or such providers to train or improve their general-purpose AI models." That covers tl;dv's own company and the AI providers it sends data to. Hosting is mainly EU, though AI features may process data in the US or the EU depending on what you select.

Otter also discloses device, cookie and location data to advertising and analytics partners. Its policy says this "may broadly be considered a 'sale' of Personal Information" under US state privacy laws. That covers product telemetry rather than meeting content, and Otter offers a "Do Not Sell or Share My Personal Information" link to opt out of it. The training use is listed under consent or legitimate interests, with no separate opt-out described for training itself.

All four re-read 24 Sep 2026. Fireflies policy updated 6 Mar 2026. tl;dv updated 9 Sep 2026. Granola effective 8 Sep 2026. Otter effective 16 Jun 2026.

Now look at the third column instead of the second one.

Two of those tools say they don't train on your meeting at all. Two train on de-identified versions of it. But all four still have to receive it, store it, and hold the information around it. That means who was invited, what the meeting was called, what's in your calendar. In one case it includes what was on your screen. In another it includes the characteristics of your voice.

A frosted container holding participants, invitees, meeting title, calendar entry, contacts, device and IP, screen captures and voice characteristics. The transcript is the smallest item in it.
Not every tool collects all of these. This is the union of what the four policies we read describe collecting. Each tool's own list is in the table above.

This isn't a criticism of how they're built. It's what their design requires, and we understand that constraint from the inside. A tool that records in the cloud has to hold the recording.

One word is worth asking about, because two of them rely on it: "de-identified."

It carries a lot of weight in Otter's and Granola's positions, and neither company publishes the method. How the identifiers are removed, and how hard the result is to link back to a person, is a fair question for any vendor whose promise depends on that word.

Who else each tool sends your meeting to

All four tools pass your meeting to outside companies for transcription or AI. Here's who each one names, and what it says those companies can do with your data.

ToolOutside companies it names for AI or transcriptionWhat it says those companies can do with your data
GranolaDeepgram and AssemblyAI for transcription. OpenAI and Anthropic for summaries. Given as examples, not a full list"We do not allow third parties (like OpenAI or Anthropic) to use your data to train their AI models"
OtterAnthropic for AI features. OpenAI to check its own models' output. Research Transcriptions to label training and evaluation dataAnthropic and OpenAI: no customer data used for training, and none stored on their platforms
FirefliesNone named in its privacy policy. The list sits in its trust centerVendors are contractually barred from training on it, and don't store meeting content after processing
tl;dvAssemblyAI and ElevenLabs for transcription. Anthropic and Google Vertex for AINot used by tl;dv or those providers to train general-purpose AI models

Granola security page(opens in new tab), read 24 Sep 2026. Fireflies privacy policy(opens in new tab), updated 6 Mar 2026, read 24 Sep 2026. Otter sub-processor list(opens in new tab), updated 31 Mar 2026, read 24 Sep 2026. tl;dv privacy policy, updated 9 Sep 2026, read 24 Sep 2026.

All four put limits on those companies in writing, and that's to their credit.

But the contract is between the note-taker and its supplier. You aren't part of it, you usually can't read it, and the list of suppliers can change after you sign up.

So a vendor review needs more than one question. It needs six.

  • What else do you hold, besides the transcript?
  • How long do you hold it, and does anyone read a sample?
  • Where does it physically sit, and under which country's law?
  • Which other companies handle it, and can I see the list? (The industry term is sub-processors.)
  • Is my protection a contract term or a dashboard setting?
  • Which of your privacy policies covers my staff?

What this means if you use Meetily

We build a meeting note-taker, so we have an obvious stake in "keep it on your own machine." Here's our answer, held to the same standard as everyone else's, including the parts that don't flatter us.

Recording and transcription run on your device by default, in both Community and Pro. A Whisper.cpp or Parakeet model does the work locally. There's no account on our side holding your audio, because there's no upload. We can't hand over what we never collected.

We also understand the pull in the other direction, because our own users felt it.

Meetily has had a pluggable model interface from day one. Local or cloud, your choice. What we didn't predict was which way people would go.

Transcription held up locally. Whisper is genuinely good. Summaries were harder. Models small enough to run on an ordinary laptop hallucinated, dropped action items, and lost track of who said what across an hour of transcript. Power users who could run both options ended up preferring a cloud model for summaries, while keeping transcription local. We wrote about that at the time in Our Quest for Meeting Summary Accuracy.

That trade-off sits underneath this whole post. The path that produces the best summaries for many people is the same path where transcript text leaves your machine.

Here's what that looks like in practice.

What you chooseWhat leaves your machineWhich companies handle it
Built-in on-device AI, or OllamaNothing. No key, no network call, no providerNone
Your own API keyTranscript text, to the provider you chose. Not audioThe provider you chose, and the suppliers it uses, under terms you signed yourself
Pro cloud transcription, off unless you turn it onAudio, to the provider you choseThe same: your provider and its suppliers, under your terms

That third row is the one exception, so here it is plainly. Pro has an optional cloud transcription path using Groq, OpenAI, or any OpenAI-compatible endpoint. If you switch it on, your audio is uploaded to the provider you picked. It stays off unless you turn it on. We'd rather write that sentence than write "your audio never leaves your device" and be wrong for everyone who turned it on.

Read that second row against the main table. If you point Meetily at your own OpenAI key, your API terms apply. That's the API column, not the consumer one. If you point it at Groq, Groq's Services Agreement applies, and that agreement bars training without your instruction.

You're not inheriting our contract with anyone, because we don't have one. You're using yours. Keep everything local and no outside company sees your meeting. Add your own key and it goes to the company you picked. That company has suppliers too, but you signed its terms yourself, so you can read them.

Bring-your-own-key is in both editions, with the same providers and the same models. It isn't a paid upgrade.

Two more things, so the tables don't oversell us. We publish write-ups for 16 providers, but the summary picker offers six: built-in on-device AI, Claude, OpenAI, Groq, OpenRouter and Ollama, plus any OpenAI-compatible endpoint. And we're not the only privacy-first option. Self-hosted stacks exist, and for plenty of teams a cloud tool with a signed DPA is the right call.

Do we train on your meetings? No. The reason matters more than the answer. It isn't a promise we're keeping. There's no path for your recording to reach us in the first place.

A privacy policy is a promise about data a company holds. The alternative is not holding it.

That applies to the next section as well.

How to check any of this yourself

Every answer above has a shelf life. Ours did, which is why we reviewed all 21 documents again in September instead of trusting what we'd written in June.

Here's the check. It takes about twenty minutes per vendor.

  1. Find the document that binds them, not the one that reassures you. Marketing pages describe intentions. Terms of service, data-processing agreements and enterprise commitments create obligations. If a claim only appears on a blog post, treat it as an intention.
  2. Check which policy applies to you. Many companies publish several by region. Find yours, and find which legal entity is named in it.
  3. Work out which tier you're actually on. Free or paid, consumer or API, personal or organization account. Then read that tier's section. This is where most reviews go wrong.
  4. Ask about training and retention separately. "We don't train on it" says nothing about how long they keep it, or whether anyone reads a sample.
  5. Ask what else gets collected, and who else gets it. Participants, calendar, contacts, screenshots, device and location, voice characteristics. The transcript is rarely all of it. Then ask for the list of outside companies that process it, and read the terms of the top two.
  6. Find out whether your protection is a contract term or a dashboard setting.
  7. Write down the date and the URL. In six months you'll want to know whether something changed or whether you misread it the first time.
  8. Repeat the check and compare.

If you want a longer version aimed at a whole vendor rather than one policy, we wrote up five questions to ask any AI meeting assistant, including us, in Where Does Your Meeting Data Actually Go? That post covers the method. This one covers the current answers.

Questions people ask

Frequently Asked Questions

Not on anything you send through the API. That's been the default since March 2023. It does train on consumer ChatGPT conversations on Free, Plus and Pro personal accounts, unless you switch off "Improve the model for everyone" under Settings, then Data Controls. Switching it off applies to new conversations, and you keep your chat history either way. Business, Enterprise and Edu accounts aren't trained on by default. Which privacy policy covers you also depends on your region, because OpenAI publishes separate Europe, US and rest-of-world versions with different legal entities named.
Often, and usually in the same direction. The API is stricter. OpenAI, Google, Mistral, Perplexity and xAI all train on at least one free or consumer tier while saying they don't train on paid API traffic. Approving a provider on its API terms doesn't approve the app your team is actually opening.
Usually, for a period. OpenAI holds API traffic for 30 days for abuse monitoring, unless you have zero data retention. Anthropic's consumer default is 30 days, and up to five years if you turn on model improvement. The Gemini app deletes activity at 18 months, but keeps human-reviewed chats for up to three years. Together AI stores prompts and responses by default while not training on them. DeepSeek publishes no deletion timeline at all.
It's split. Fireflies and tl;dv both say customer content isn't used to train AI models. Granola trains only on de-identified data and lets you opt out. Otter's policy says it does use de-identified recordings and transcriptions to train its own AI. For all of them, the bigger question is collection rather than training, because a cloud note-taker has to receive and store your recording, and usually your participant list and calendar with it. Most also pass your meeting to outside transcription and AI companies, under contracts between them and those companies.
Transcript text, and only if you've set up bring-your-own-key summaries. Audio stays on your device unless you've deliberately switched on Pro's optional cloud transcription. If you use the built-in on-device model or Ollama, nothing is sent at all.
No. Your recordings and transcripts stay on your machine and we never receive them, so there's nothing for us to train on.
We're not going to rank them. We sell one of the options on this page, so a ranking from us would be worth less than the terms themselves. The tables are here so you can check the tier you're actually on against the promise you think you have.

Everything we've said about another company comes from that company's own published documentation, reviewed between 16 and 24 September 2026. Published statements aren't audit findings, and these documents change without notice. Check again before relying on any of this for a compliance decision. If we have a policy wrong or out of date, tell us and we'll correct it and say that we did.

Try it: Meetily Community is free and open source under the MIT license. Pro has a 14-day free trial, no payment needed. The per-provider references are at LLM Privacy.

Get started

Ready to try Meetily?

Join 560,000+ users who use Meetily for private meeting transcription. No bots, privacy first. Community Edition free.

No meeting bots
100% local transcription
Free & open source
Download Free

Star on GitHub (30K+) · Open source & self-hostable

Get Started with Meetily

Meetily Pro

Advanced features for individuals and teams.

Download

Get Meetily for Mac or Windows. Free and open source.

Download

Recent Articles