AI writing still has telltale signs, study finds, despite efforts to sound human

Graphite research identifies thousands of phrases that give away machine-generated prose

By LineZotpaper
Published
Read Time2 min
Sources2 outlets
A new study from marketing firm Graphite finds that frontier AI models still produce writing with detectable quirks, even as they eliminate older tells like em-dashes. The research identifies thousands of phrases that appear far more often in AI-generated text than in human writing.

Researchers at marketing firm Graphite analyzed the writing habits of leading AI models, looking for words and phrases that betray machine authorship. The study, which used a corpus of 10,000 articles published before ChatGPT as a human control group, found 13,000 phrases that were at least twice as common in AI content as in human writing.

Each model shows its own signature. Claude Opus 5.5's biggest tell is the word "dependable," which appears 23 times more often than in human samples. It also favors constructions like "more than an X, it's a Y" and phrases such as "this matters" (116 times more often) and "why X matters" (92 times more often).

OpenAI's Astra model, meanwhile, often describes things as having "another dimension," hedges claims with "may provide" or "can provide," and uses what Graphite calls "corrective framing," such as "not simply X" or "rather than relying on X." These constructions were more than 100 times more common in Astra-generated prose.

Notably, all frontier labs appear to have responded to criticism about overusing em-dashes. Graphite found that Opus 5.5 uses the punctuation mark 99% less often than Opus 5, Astra uses it 88% less than humans, and Gemini 3.1 Pro has nearly eliminated it.

Despite these changes, the overall number of tells is holding steady. "It's not like the tells are decreasing," Graphite's chief AI officer Greg Druck told TechCrunch. "They are managing to remove the most well-known tells, but other ones pop up. And every model version has its own."

The persistence of these patterns is notable given the emphasis on natural writing in recent model releases. Anthropic, for example, touted Opus 5.5's ability to "communicate more naturally than prior models," and OpenAI made similar claims for its latest models. The study suggests that while models can drop specific habits, they develop new ones.

§

Analysis

Why This Matters

  • As AI-generated text becomes widespread, the ability to detect it has practical implications for journalism, academia, marketing, and content moderation.
  • The study shows that AI models have not achieved truly human-like writing, which could affect trust in automated content and the credibility of AI-assisted tools.
  • Understanding current tells may inform future detection methods and push labs to further refine their models.

Background

Graphite is a marketing firm that studies AI-generated content. The company's latest research updates earlier work on AI writing tells, such as the overuse of em-dashes and words like "delve." The study design involved having AI models rewrite pre-ChatGPT articles from summaries to minimize source bias, then comparing word and phrase frequencies with human-written text. This methodology aims to isolate model-specific habits rather than stylistic echoes from the source material.

Key Perspectives

AI Labs (Anthropic, OpenAI): They claim their models communicate more naturally and clearly than before, emphasizing improvements in fluency and readability. The persistence of unique tells, however, suggests that these efforts have not fully succeeded. Graphite (Researchers): The firm argues that AI-generated writing remains distinguishable at scale. Their data shows that while models remove known tells, new ones emerge, making detection an ongoing challenge. Critics/Skeptics: Some may question whether frequency-based tells are reliable across different contexts or user prompts. Others might argue that as models improve, the gap between AI and human writing will narrow, making such detection less useful over time.

What to Watch

  • How AI labs respond to the study: whether they attempt to purge these new tells in future model versions.
  • Whether detection tools incorporate Graphite's findings to improve AI-content identification.
  • The evolution of tells as models are updated: whether the overall count eventually declines or continues to fluctuate.

Sources

Zotpaper

Written by software from the reporting listed above, scored by an automated standards desk, and published without a person reading it first. If something here is wrong, tell the editor and it will be put right.