Saltar al contenido

How AI fingerprinting works: the real story of how they try to detect your texts (and why they fail)

How AI detector works: the story of fingerprinting that tries to decide if you wrote your own text

Cómo funciona un detector de IA

There is a scene that repeats every day.
A student submits an assignment. A journalist publishes a report. An editor uploads an article. Everything seems normal… until a message appears on screen:

“Probability of AI-generated content: 87%”

Without explanation. Without context. Without right to appeal.
As if an algorithm had read your text and, in its infinite statistical wisdom, had decided that you are not you.

This is how the new fever for detecting AI works. A mix of haste, fear, ignorance and blind trust in tools that do not understand language, but act as if they could read the writer’s soul.
This TechLab is the story of how we got here… and why the system fails.

1. The origin of the problem: wanting to measure what cannot be measured

Everything starts with a question that seemed innocent:

“Can we distinguish a human text from one created by AI?”

Scientifically, the correct answer is: no, at least not in a reliable way.
But the industry wanted to try anyway. And thus was born “fingerprinting”, a set of techniques that promised to identify hidden authorship in any text.

It sounds sophisticated, almost detective-like. But it is enough to open the lid to see gears more fragile than they want us to believe.

2. Perplexity: when writing well turns you into a suspect

The first technique adopted by detectors was to analyze how “predictable” a text is.
This is called perplexity, a metric that sounds like quantum physics but behaves like a deaf tracking dog.

Imagine standing in front of a detector that evaluates you thus:

  • If you write with clear phrases, well ordered, without sharp turns → AI.
  • If irregular phrases appear to you, ideas that clash with each other, some loose error → human.

The detector does not seek intelligence: it seeks noise.
That natural noise of human writing that disappears when you correct your text or edit it with an AI.

That is why it is not strange that a worked article marks as AI while a mediocre text passes as human without discussion.

3. Burstiness: the rhythm that decides if you “think too well”

The second technique is burstiness, which analyzes the rhythm of the text.
Humans, when writing, accelerate and brake without realizing it.
But AI models, until 2022, maintained an almost perfect rhythm.

What is the problem?
In 2025 models are capable of imitating human chaos on demand.
You can ask them for short phrases, long ones, disordered, with ups and downs in tone.
And they do it better than many humans.

So the “rhythm test” now serves to detect… nothing concrete.

4. Watermarking: the watermark that disappears with just breathing

There was a moment when the community believed they had found the definitive solution: placing invisible marks (watermarks) in texts generated by AI.

An elegant idea:
If the AI leaves a mathematical trail, we will be able to discover if it generated the text.

But here comes the twist:

  • if you translate the text → broken watermark
  • if you summarize it → broken
  • if you reorder ideas → broken
  • if you mix it with human text → broken
  • if you pass it through a model that does not use watermark → non-existent

This fragility is documented in studies like this.

And to finish:

most open models (Llama, Mistral, GPT-J, GPT-OSS) DO NOT carry watermark.

It is like trying to identify criminals by their perfume… when half use cologne and the other half did not shower.

This method is not theoretical: it was described in detail in the study of Kirchenbauer et al. (2023), where it is proposed to divide the vocabulary into “green” and “red” tokens.

5. Stylistic fingerprinting: searching for traces where there are none

The third technique attempts to identify stylistic patterns.
A subtle movement, almost literary: analyze the way you connect ideas, how you use adjectives, the structure of your sentences, even how a model tokenizes your words.

On paper it seems brilliant.
Until you remember two details:

  1. Humans do not have a unique style.
  2. AI can imitate any style if you ask it.

This turns the task into an infinite game of cat and mouse.
An AI detector trained on GPT-3.5 cannot recognize either GPT-4 or Llama 3.
And much less open models, fine-tuned, modified or mixed by the community.

The result:
a detector that tries to recognize shadows moving on the wall.

6. The hidden tragedy: false positives

Here we arrive at the most painful point.

The majority of the damage these detectors do is not failing to identify AI.
It is failing to identify humans.

A study published in Education Integrity confirms this pattern of false positives in real educational contexts.

Who are the most punished?

  • people who write very well,
  • non-native students,
  • editors who follow SEO guidelines,
  • journalists who correct style,
  • and anyone who uses AI as editor, not as author.

And it is that when an AI corrects a human text, it softens its natural irregularity.
That is enough for the detector to classify it as generative AI.

Of course:
it does not know how to distinguish editing from authorship.

It believes a polished text is “too perfect for a human”.

An offensive idea, but real in the world of detectors.

7. The impact of open models: chaos in the fingerprint registry

So far, detectors tried to follow the rhythm of major commercial models.
But came the OSS revolution: Llama, Mistral, GPT-J and an army of variants fine-tuned at home, mixed, cut, adapted to niches and launched without watermark.

Fingerprinting collapsed.
There is no “universal signature” that these models share.
Each one writes as it wants, depending on who trained it or how.

It is like trying to recognize an artist by their brushstroke… but now millions of artists share the same brush.

8. The blind spot we should fear: style surveillance

Beyond the technical, there is a social issue rarely mentioned:

if your style can be used to judge you, then writing becomes a monitored act.

Content is no longer evaluated, but its “probabilistic origin”.
The polished text is penalized.
Suspicion falls on the one who writes with clarity.
And justification of authorship is forced as if every phrase were a trial testimony.

Instead of protecting academic integrity, detectors are creating an atmosphere of distrust about language.

And that should indeed worry us.

The detector does not detect AI, detects patterns… and gets confused by them

After reviewing perplexity, burstiness, watermarking, stylometry and the chaos of the modern ecosystem, the conclusion is as simple as it is devastating:

An AI detector does not identify authorship; it classifies probabilities.
It does not measure if you wrote the text; it measures if your style resembles that of an AI.
And in 2025, writing well looks too much like writing “like an AI”.

Detectors generate false security, false accusations and a growing sense of surveillance.
They are not a neutral tool.
And they should not dictate sentences on human works.

OpenAI officially removed its AI Text Classifier in 2023 due to low precision and high false positive rate.

Even US universities recommend not using detectors as conclusive proof, as explained by Illinois State University.

Language deserves something better than a digital traffic light.