The re-identification problem that breaks "anonymized" data

"Anonymized" is one of the most misleading words in data privacy.
It sounds like a guarantee: this data has been stripped of anything that could identify you. In practice, it is closer to a description of a process that was applied, and whether that process actually protects you depends on assumptions about the world that often turn out to be wrong.
The re-identification problem is not a niche academic concern. It has surfaced in medical records, location data, movie ratings and genomic databases. Understanding why it happens is useful for anyone building systems that collect, process or publish data about people.
What anonymization is actually trying to do
Anonymization is the process of transforming a dataset so that individuals cannot be identified from it. The goal is to make it safe to publish, share or analyze data without exposing the people the data is about.
The naive version is simple: remove the obvious identifiers. Name, address, phone number, email, strip those fields and the data is anonymous.
This works until someone has another dataset.
The problem is that most data that is interesting to collect is also linked to other data that exists in the world. Your medical records without your name are still linked to your age, your postcode, your employer, your diagnosis. If someone has access to any other dataset that contains some of those attributes, they can potentially match your record to you.
This is the core re-identification insight: anonymization is not a property of a dataset. It is a property of a dataset in relation to all other data that exists.
k-anonymity: a formal attempt at a guarantee
k-anonymity is the most widely used formal model for measuring anonymization quality. A dataset satisfies k-anonymity if every record is indistinguishable from at least k-1 other records with respect to the quasi-identifiers, the attributes that could potentially be used for re-identification.
If a dataset is 3-anonymous, every combination of quasi-identifier values appears at least 3 times. An attacker who knows your age, postcode and sex cannot narrow you down to a unique individual, there are at least 2 other people with the same combination.
k-anonymity is a genuine improvement over naive anonymization. It gives you a formal bound: even with access to your quasi-identifiers, an attacker faces at least k candidates.
The problem is what k-anonymity does not protect against.
Homogeneity attack: if everyone in a k-anonymous group has the same sensitive value, the same diagnosis, the same salary band, knowing which group you are in tells the attacker your sensitive value, even if they cannot identify you by name.
Background knowledge attack: an attacker who knows something about you that is not in the quasi-identifiers can use that knowledge to narrow down which record is yours, even in a k-anonymous dataset.
The curse of dimensionality: as the number of attributes in a dataset increases, k-anonymity becomes harder to achieve without destroying data utility. With enough attributes, nearly every record is unique, and making every combination appear k times requires suppressing or generalizing so much data that it becomes useless.
The cases that made this concrete
The Latanya Sweeney re-identification of Massachusetts health records (1997)
When Massachusetts released "anonymized" health records for research, they removed names and social security numbers. Latanya Sweeney, then a graduate student at MIT, matched the released records to a public voter registration database using just three fields: date of birth, sex and postcode. She was able to identify the medical record of the state governor.
She later estimated that 87% of Americans could be uniquely identified by the combination of date of birth, sex and five-digit postcode alone. The naive removal of name and SSN was not sufficient.
The Netflix prize re-identification (2008)
Netflix released a dataset of 100 million movie ratings for a machine learning competition, removing usernames and replacing them with random IDs. Arvind Narayanan and Vitaly Shmatikoff showed that by comparing the anonymized Netflix ratings to public IMDb reviews, which include both ratings and dates, they could re-identify individual Netflix users with high confidence.
The attack required only a handful of movies rated in common. Sparse but distinctive patterns in behavior are surprisingly identifiable even without any obvious identifiers.
Location data and the uniqueness of movement (2013)
A study in Nature showed that 95% of individuals in a mobile phone dataset could be uniquely identified by just four location points, four places they had been at four times. Location data is often treated as less sensitive than medical or financial data, but individual movement patterns are highly distinctive. Two people rarely go to the same four places at the same four times.
Genomic data re-identification
Genomic data is perhaps the hardest case for anonymization. DNA is by definition a unique identifier. Studies have shown that individuals can be re-identified from supposedly anonymous genomic datasets by cross-referencing with public genealogy databases, recreational DNA testing results or other genomic datasets. Removing a name from a genome sequence does not make it anonymous.
A concrete example of how a biometric system can be designed with the re-identification problem in mind is this overview of proof of human and data handling, which covers what is retained after verification and why.
What follows from this
The re-identification literature establishes a few things that should inform how privacy protections are designed.
Anonymization is not binary. A dataset is not anonymous or not anonymous, it is more or less re-identifiable depending on what other data exists in the world, who has access to it, and what computational resources they can bring to bear.
The threat model has to be realistic. k-anonymity and similar models assume a specific adversary with specific capabilities. As the availability of auxiliary datasets has grown, census records, voter rolls, social media profiles, commercial data brokers, the realistic adversary has become significantly more capable than the models assumed.
Utility and privacy are in genuine tension. The more you generalize and suppress data to achieve formal anonymization guarantees, the less useful the data is for its intended purpose. This tension cannot be fully resolved, it can only be managed.
Differential privacy is the current state-of-the-art formal approach. Rather than trying to make individual records unidentifiable, it adds carefully calibrated mathematical noise to query results, providing a formal bound on how much any individual's data can affect the output. It is more principled than k-anonymity but comes with its own trade-offs in terms of utility and complexity.
Why this matters for system design
If you are building a system that collects personal data and plans to publish, share or analyze it in anonymized form, the re-identification literature suggests a few practical things.
Do not rely on removing obvious identifiers. The fields that feel obviously identifying, name, SSN, are rarely the only re-identification vectors. Quasi-identifiers including age, location, timestamps and behavioral patterns can be as dangerous in combination.
Think adversarially about auxiliary data. The question is not "can someone re-identify this dataset in isolation?" It is "can someone re-identify this dataset given everything else they might have access to?" The answer depends on what other datasets exist and who might try.
Consider whether you need to publish the data at all. Differential privacy, federated learning and other privacy-preserving computation techniques can often answer the questions you actually need answered without publishing the underlying data. If the goal is to learn something about a population, you may not need individual-level records.
Be honest about what anonymization actually guarantees. Telling users their data has been anonymized, without specifying the threat model that guarantee holds against, is at best imprecise and at worst misleading.
The gap between "anonymized" as a label and "anonymous" as a property is where most privacy failures happen.
Related reading





