When it comes to AI tool usage, the “Not Used for Training” data promise does not mean you're out of the woods

When it comes to data retention and training on data itself, they often mean two different things within the privacy policies of closed-source model providers and AI tools that leverage these closed-source models.

By Wellington Johnson, Co-Founder··7 min read·Updated August 13, 2026
Closed-source modelsData PrivacyData Retention

This article is educational and not legal, clinical, or compliance advice.

One of the most important things that gets looked over as closed source large language models, or LLMs, are being used are data privacy policies. They're disregarded just as often as Terms of Service statements, sometimes more because we rarely care enough about how our data is being used. Data brokers have been around almost as long as the internet hit global scale, and the movement and storage of data has become even more important in the age of AI.

Closed source model companies like OpenAI and Anthropic train their models on trillions upon trillions of bytes of data scattered across the internet containing information about you, me, and countless people across the globe. And while that's allowed us to have and use this incredible technology, that also means that these LLMs can really only get generally better intelligence-wise with more of that data being passed into them and trained on. That's fine for many people who don't think much of pasting their budgets, resumes, and workout goals into a chatbot, but it might not be fine if you're using these tools with sensitive info. For many professionals out there, confusion is still prevalent when trying to understand how they should go about using AI in their day to day tasks while still maintaining a semblance of privacy when it comes to proprietary data, client information, and otherwise sensitive details they wouldn't want retained in perpetuity or trained on.

I'm going to talk through why data privacy/retention language that points to data not being trained on and data not being retained are separate and often muddy, and what can be done to still use those tools effectively regardless of the policies they have.

First, I'd like to set a ground rule that is the basis of this entire article, which is: Just because a company says they don't train on your data, doesn't mean they don't keep your data for a long time. Pay attention to mentions of both data retention and data training in these policies.

So why is training on data and data retention often different concepts in these policies?

These terms cover different parts of the data lifecycle. Retention is about what's in possession, whether a company keeps your data and for how long. Training is about the usage of that retained data, whether that data is fed into a process that improves the model or tool. When a company says it doesn't train on your data, it's making a specific promise about how that data will be used, but it isn't necessarily promising that the data won't be stored. In the same way, a short retention period doesn't explain everything that may happen to your data before it's deleted.

Privacy policies can get really weird because training and retention are often presented very close together, which makes them sound interchangeable even though they aren't. You have to read each promise separately to understand what you're actually agreeing to.

So what happens when a model is "trained" on your data?

Without going too deep into generative AI training practices, your data can be used for training in two phases broadly, the first phase being in a "pre-training" phase where patterns are learned from your and many other individuals and companies questions, responses, documents, etc and by an algorithm where it's main goal is to predict the next word given a word or series of words. The second phase is "post training", where the model goes from an unwieldy next word predictor to a more refined word predictor that can follow instructions well, not give harmful feedback, basically it's a more chill version of it's pre-trained self.

Both of these phases often need a considerable amount of data in order to product a useful model, especially one that is demonstrably better than its predecessor. Going from Anthropic's Opus 4 to Opus 5, no matter how you feel about the quality, took a ton of data to ensure that it can reason that much better over a broad set of problems that a human might throw at it. Same with OpenAIs GPT models, all of them need. your. data.

Okay, so if a company doesn't train on my data, why should I care about retention?

1. Security Breaches

The thing about data retention is that, just because it's not being used for anything, doesn't mean it's not valuable to other people outside of that company that's not using it. With the rise of AI, individuals and companies alike are able to build software within hours and days that previously took months and sometimes years. But with that benefits comes drawbacks, individuals and groups of people are now able to create exploits and hack faster than they ever have before, and in many more interesting ways.

It makes sense then that data breaches could reasonably become more commonplace than they already are, and if your data is being stored somewhere that's not taking advantage of AI to defend itself as comprehensively as potential malicious actors are to build new weapons to attack these companies and individuals, then it seems like not a great idea to just have sensitive data sitting there.

2.  Other internal processes that leverage your data

A lot of data retention policies center around needing your data for four things:

1. Abuse Monitoring - Companies need to be able to detect and prevent harmful activities like generating hate speech, facilitating cyberattacks, or producing illegal content.
2. Debugging - If a tool crashes or acts up, having the logs for those incidents might contain your data, and other data that's retained to corroborate the story will also be used.
3. Legal Compliance - A lot of times companies are required by law to retain certain data to comply with laws or respond to court orders and defend against lawsuits.
4. Quality Assurance - Internal QA processes will have access to your data at certain times to help with solving customer issues, as well as label data to grade an AI tool's helpfulness.

All of these reasons require having your data and leveraging it in ways you might not feel comfortable with, especially if it contains sensitive information.

So how can I protect my sensitive data while still using these tools and models?

Honestly, the most practical way to protect sensitive information is to just remove it before it even reaches an AI tool. Privacy policies can change and shifting retention windows can be hard to understand, so data that isn’t used for training can still be used for other reasons. At that point you’re relying on someone else’s policies and security practices to keep it safe.

That’s where Nonymize comes in. Nonymize identifies and replaces personal or confidential details in sensitive conversations while preserving the context that makes the material useful. You can then use that anonymized version with the AI tools you already rely on without unnecessarily exposing client names, company information, financial details, or other identifying data that you're unsure or uneasy about sharing.

No solution removes every possible risk, and it’s still worth understanding the policies of the tools you use, but Nonymize gives you control over the most important part of the process: what information leaves your hands in the first place.

Check us out!

Protect your transcripts before AI analysis.

Create an account to prepare sensitive transcripts before they reach AI.

Sign up