Anonymization Is More Than Removing a Person's Name

Removing names from a client transcript isn't anonymization — indirect identifiers like employer, job title, city, and deal details can re-identify someone with no name in the document at all. Research from Carnegie Mellon found 87% of Americans are uniquely identifiable from just ZIP code, birth date, and gender, and a peer-reviewed 2019 study put re-identification at 99.98% with 15 attributes, even in incomplete datasets. Real anonymization means catching the full combination of identifiers, keeping replacements consistent across documents, and putting a human review step in front of anything that reaches an AI tool.

By Nonymize Editorial Team··7 min read·Updated September 2, 2026
AI privacyClient confidentialityData anonymization

This article is educational and not legal, clinical, or compliance advice.

Anonymization Is More Than Removing a Person's Name


A transcript can be nameless and still identify anyone in any room everywhere. A document can be void of address and phone number and still break privilege. 

Sit down, we need to have a talk. In today’s world, in 2026, you don’t need someone’s name to know who they are. Here is a fact: With three or four details that are specific to that person, you have enough information to find out who it is. 

And in your weekly client calls? Your monthly check-ins? Those weekly team huddles? In any real conversation your meeting notetaker is recording, those 3-4 details are everywhere.

Why "I deleted the names" feels like enough

The instinct isn't stupid, frankly at first pass it logically makes sense. It's how we think about anonymity everywhere else - a masked face, an unsigned letter, a blurred photo. Names are how we index people. Remove the name and the person becomes ‘unfindable’.

That's how anonymity works in the physical world. It is not how identification works in the digital world.

Two kinds of identifiers, and the one you miss might be more damaging

There are two kinds of identifiers: Direct and Indirect. Direct identifiers point at one person by themselves. Name, email address, phone number, account number, Social Security number. These are the ones everybody immediately thinks of to remove when trying to protect someone. Direct identifiers are the reason we think, “well I hid her name or removed his phone number” means we’ve completely protected their identity. 

Indirect identifiers point at nobody on their own and at exactly one person at the same time. Employer. Job title. City. The conference they spoke at last month. The grant they won. The deal that closed in March. The fact that their co-founder used to work (somewhere notable). The epic family trip they went on in 1996 and was interviewed by the Today Show. 

When you delete the direct identifiers you give yourself the illusion that you’ve removed everything that could identify the person. What’s actually happened is you might’ve left the most damning of identifiers right there in the transcript or document you’re about to upload to AI.

The research is worse than you'd guess

In 2000, Latanya Sweeney of Carnegie Mellon University, released a research report that found that 87% of Americans can be uniquely identified using just three data points: ZIP code, birth date, and gender. Three fields. None of which are ‘name’.
Some 19 years later, Luc Rocher, Julien M. Hendrickx & Yves-Alexandre de Montjoye found that 99.98% of Americans would be correctly re-identified in any dataset using 15 demographic attributes. The pier-reviewed paper's whole point is that incomplete and heavily sampled datasets still re-identify. So the "but I removed some of it" defense is absolutely blown-up and it's addressed with math rather than assertion. 
Sit with that for a second. 

The findings of the latter offer an even more frightening finding - their conclusion is that even heavily sampled anonymized datasets are unlikely to satisfy the modern standards for anonymization set forth by GDPR, and that the finding seriously challenges the technical and legal adequacy of the de-identification release-and-forget model. 

What this actually looks like

The pandemic ushered in the work-from-home, meet virtually, co-work virtually era probably 10-15 years faster than it would’ve otherwise happened for ‘remote possible’ jobs. The fact that it was forced down our throats for two years accelerated a change that while yes it was coming, the reality was it was nowhere close to the adoption rate it forced, 90%+. And the residuals of the pandemic is that hybrid or remote work is not only hear to stay, but it’s now the majority for remote-possible jobs. According to a 2026 Gallop report - 52% of remote-capable jobs are hybrid and 26% are fully remote.

That’s a combined 74% of remote-possible jobs being either part-time remote or fully-remote. 

Meeting notetakers, recorded calls, they’re the new zip drive and notebook. 

So you have one of your weekly client calls. You download the transcript from your notetaker and you Command + F  find all of the names and ‘redact’ them - client, colleagues, your name, their company, all of the brands/companies you remember named. The transcript is now clean, by most people's standard. 

Except it isn’t. Not even close, actually. Here's what's still sitting in there:
The industry. The city. The size of the team. The conference they're speaking at next month. The funding round they just closed. The name of the spa they mentioned as their "self-care thing." The softball team their company sponsors. The full detail of how you both met and what you did the first time you met. The specific number they're trying to hit this quarter. The thing they're known for on LinkedIn.

No name anywhere in that document, and yet it’s a treasure trove of information putting a ‘name’ to them. 

Arm anyone in that industry with all of that information and ask them to name who it is. They’ll be able to do so in under a minute. Not by cracking anything, not by any lucky guess - just by process of elimination. There aren't that many companies in that city, that size, with a founder who spoke at that conference, who they know you know, and recently closed a funding round.

That's not a hypothetical. That's every transcript you or I've ever redacted by Command + F.

Why find-and-replace can't fix it

The problem isn't effort. It’s not even time.  It's our memory.

Find-and-replace only catches what you already remembered to search for. To do it properly, you'd have to recall every identifying detail from an hour long conversation to 100% accuracy before you ever open hit Command + F to search. 

Spoiler alert: Nobody does that. And the things you miss aren't the obvious ones. They're the nickname they use for their business partner. The company name from before the rebrand, mentioned once in passing. The street they said they were walking down when they got the news.

You're not being careless. You're being asked to pass a memory test that's designed with failure all but guaranteed.

What real anonymization has to do

Three things.
Catch the combination, not just the name. Names are but one field. Organizations, roles, locations, dates, financial figures, and project references are alllll of the rest of it - and identity lives in how they stack up, not in just one of them. 

Offer clarity and consistency. If you are wanting to upload anonymized weekly transcripts into your AI workflow, consistency matters. This is for both the same document and multiple document instances. If your client is [Person A] and you’re [Person B] and the next you’re both [Redacted] - you’re going to have a tough time knowing who’s who, and your AI is absolutely going to fail at it. When the same person becomes a different placeholder every time they're mentioned, the document stops making sense and you've traded one problem for another. 

Same entity, same replacement, throughout.

Let a human check it. No detection system catches everything - not Nonymize, not anyone's. The person who was in the room is the only one who knows what's missing from the page. That's not a limitation to apologize for. It's the reason the review step exists. That’s why we made humans giving the thumbs up a requirement in our product - you will always have final say.

The question to actually ask

"Did I remove the names?" is the wrong question to be asking. "Could someone who knows this industry figure out who this is?" - If the answer is no, then you’ve anonymized effectively. That's a harder question, and almost every single find-replace attempt fails it…badly.
 
It's also the question we built Nonymize to answer - catching all the combinations, keeping it consistent, and giving the human the final say before anything goes anywhere.


Protect your transcripts before AI analysis.

Create an account to prepare sensitive transcripts before they reach AI.

Sign up