As AI-generated prose spreads across the web, readers are hungry for ways to spot it. The early tells — a fondness for em-dashes, the word "delve" — have largely been scrubbed away. But a new study from the marketing firm Graphite finds that frontier models still leave plenty of fingerprints, and that each one has its own.
To measure them cleanly, Graphite started with a corpus of 10,000 articles published before ChatGPT's release as a human control group. It then had different AI models rewrite those articles from summaries, stripping away the original phrasing so any leftover patterns would come from the models rather than the sources. With matched samples, the researchers could compare how often particular words and phrases appeared, and examine broader sentence construction.
The result: 13,000 phrases that turned up at least twice as often in AI text as in human writing — Graphite's definition of a "tell." The scope surprised the researchers. "It turns out that Claude models are actually getting closer to the human word distribution over time," Graphite's chief AI officer Greg Druck told TechCrunch. "And for the GPT models, it's getting further away."
Claude Opus 5.5 has a distinct profile. Its single biggest giveaway is the word "dependable," which appears 23 times more often than in human samples. The model has largely shed the old "it's not X, it's Y" construction, but still favors a related move — saying something "is more than an X, it's a Y." Above all, it likes to explain significance: "this matters" showed up 116 times more often than in human writing, and "why X matters" 92 times more often.
OpenAI's Astra reads differently. It frequently describes "another dimension" of a topic and hedges claims with verbs such as "may provide" or "can provide." Its signature, per Graphite, is "corrective framing" — defining a topic by what it is not, using constructions such as "not simply X" or "rather than relying on X," which appeared more than 100 times as often as in the human control set.
One tell has been nearly eliminated. Under pressure from years of mockery, the labs have all but retired the em-dash: Opus 5.5 uses it 99% less often than Opus 5, Astra 88% less often than human writers, and Google's Gemini 3.1 Pro has almost removed it entirely. But suppressing one marker has not reduced the overall count. "It's not like the tells are decreasing," Druck said. "They are managing to remove the most well-known tells, but other ones pop up. And every model version has its own."
That persistence sits awkwardly next to the labs' own messaging. Anthropic promoted Opus 5.5 as communicating "more naturally than prior models," and OpenAI said users of its GPT-6 Sol and Luna could expect "more clarity, less jargon, [and] fewer odd turns of phrase." Druck is skeptical that the labs can fully eliminate telltale constructions: "These are giant models with billions of parameters," he said. "They have some finite number of tests they can run, and things slip through."
One caveat matters for anyone tempted to use this as a detector: these are corpus-level statistics. They describe patterns visible across large collections of text, not the probability that any single sentence was machine-written. A person can naturally write "this matters," and a model can avoid its favorite tell entirely. What the research really shows is that AI style is a moving target — and that removing one signature simply clears space for the next.
Comments (0)
Log in to join the discussion
Log InNo comments yet