The uncomfortable truth: you can’t perfectly tell, but you can stay skeptical
AI-generated text is writing produced or heavily shaped by large language models, often polished enough to pass as human yet carrying invisible patterns, watermarks, and occasional hallucinations that betray automated authorship when you know what to look for and accept that detection is always uncertain and imperfect. To detect AI-generated text today, you must treat every document as a mix of three layers: what your own reading instincts say, what machine signals reveal, and what the underlying evidence can prove. Relying on any single layer is naïve. Watermarks help, but they are not universal; AI content detection tools can highlight suspicious passages, but they produce false positives and can’t reliably prove AI involvement either way. The only safe stance is critical, source-first reading: assume the words might be automated and ask whether the claims are real, sourced, and coherent.
How AI watermarks text and what that means for detection
The most important shift you need to grasp is that AI systems now watermark raw text, not just images. Companies have found a way to watermark raw text too, and they’re doing it by making the text itself, or rather the order that text is being generated in part of the watermark. To understand how watermarking AI-generated text works, remember that a model writes by predicting tokens – small units that can be whole words, parts of words, or punctuation – one at a time based on what is most likely to come next. When two tokens are equally plausible, companies adjust their models to favor one identical score over another, using a hidden pattern tied to earlier words, slightly boosting one option and lowering the other. The exact pattern of where these nudges happened is the actual watermark. Companies have their own detectors, like Google’s SynthID, that know exactly what happened and instantly recognize the pattern. This turns provenance into a technical fingerprint instead of a visual label and quietly changes how AI content detection will work going forward.

Why receipts and machine marks beat old-school AI content detection
Most public tools that claim to detect AI-generated text are unreliable. They lean on surface patterns – sentence length, word choice, repetitiveness – and therefore misfire on polished human writing while missing carefully edited AI prose. They produce false positives and can’t reliably prove AI involvement either way. That is why watermark-style receipts matter more than stylistic guessing. Anthropic is working to include machine-readable marks in content that Claude generates to support transparency and follow legal obligations. Starting with Claude models launched on or after August 2, 2026, generated content will carry machine-readable provenance marks from the moment the model produces it. Content now carries a machine-readable signal identifying that help was involved. As AI-generated content becomes commonplace, greater transparency and signals about where content comes from can give people useful context about the information they consume. This approach is far more honest than pretending detectors can unmask every AI paragraph; it tells you when assistance was used instead of guessing from vibes.

PwC’s fake citations mess: how to spot hallucinated references
If you want a real-world lesson in why you must spot fake citations, look at what happened when a major firm published AI-shaped reports as thought leadership. Several reports included AI hallucinations, fabricated citations, and fake footnotes that were later identified as part of a pattern of irresponsible AI usage. Vibe citations are hallucinated references to works that an AI-generated document cites as genuine. They can be demonstrated as fake due to issues such as non-existence, incorrect authors, incorrect publication dates, erroneous URLs, or the citation may simply not exist at all. To spot fake citations in any suspicious document, treat references as claims, not decorations. Pick a few at random and check whether the work exists, whether the authors and dates match, and whether the link leads to what is promised. If a paper is full of broken or non-existent references, you are not looking at serious research; you are looking at AI-generated text with evidence glued on for show.

The arms race: humanization tools and the limits of watermarking
As soon as AI watermarks text, other tools try to strip or hide those marks. Heavy editing, paraphrasing, translation, or mixing Claude’s output with other writing can alter the text enough that the watermark is undetectable. You can remove the watermarks (sort of). That means any policy that treats watermarked text as the only proof of AI help – or its absence as proof of human purity – is already behind the curve. Also, if you try to play the system and ask an AI to randomize a text generated by another AI to remove the watermark, it would probably still apply its own watermark. Humanization tools exist largely to make AI-written paragraphs look less like AI and to evade AI content detection, which creates a pointless arms race between detectors and evaders. Instead of chasing perfect detection, institutions should focus on evidence quality, source transparency, and accountability for the claims being made. Your goal as a reader is not to be a forensic analyst; it is to refuse to be fooled by polished nonsense.




