Anthropic just shipped its new flagship large language model (LLM) Claude Opus 5.5 last week, and its word choices resemble human writing more closely than its predecessor’s, according to a new analysis from digital marketing agency Graphite.
But anyone hoping the model has shed its recognizable habits will find plenty left: the researchers identify 2,548 words, phrases and sentence patterns that appear at least twice as often in its generated articles as in comparable human articles.
The update to Graphite’s AI Tells study, shared with VentureBeat ahead of its planned Thursday publication, adds Anthropic’s latest Opus model to an analysis of how AI writing changes between generations. The comparison covers articles from ten models and human writers on 9,974 matched topics.
Its central finding concerns word-distribution divergence, which measures how differently two collections of text use words overall. Opus 5.5 scores 0.052 against the human sample, down from 0.064 for Opus 5, a 19% reduction. OpenAI’s newest flagship model GPT-6 Astra, released earlier this month, moves in the opposite direction, scoring 0.109, up from 0.101 for GPT-5.6 Sol.
The results offer enterprise content teams a way to examine model behavior beyond coding and agent benchmarks. They also show why evaluating AI writing requires more than checking whether a model has stopped using a handful of familiar expressions.
“I think what we see here is not that there's a steady progression of it being better or worse,” Graphite CEO Ethan Smith told VentureBeat in a video interview. “It's just sort of different, and each one is different, which is kind of surprising.”
The em-dash nearly disappears
Graphite counts just 0.015 em-dashes per 1,000 words in Opus 5.5’s articles, compared with 2.92 in Opus 5’s, a decline of about 99%.

The new model uses them much less often than the human writers in the sample. Astra also uses them well below the human rate.
That complicates the popular association between em-dashes and machine-generated copy. A punctuation mark that attracts suspicion can become uncommon in newer models while remaining part of human writers’ ordinary vocabulary.
Opus 5.5 also earns a lower score for what Anthropic calls “mannered prose,” language that favors metaphor or flourish over direct expression. Graphite’s score falls 37%, from 16.75 for Opus 5 to 10.57 for Opus 5.5. It remains above the human sample’s 6.65 and Astra’s 7.91.

Graphite uses Claude Opus 5 to assess this quality in a subset of 1,000 matched topics. The score therefore reflects a model’s judgment of style, rather than a mechanical count of ornate sentences.
Gregory Druck, Graphite’s chief AI officer, said asking a model to avoid this style has proved useful in his own writing experiments.
“I've been lately a lot telling it not to do mannered prose,” Druck said. “That's a pretty good one.”
But the study does not equate stylistic similarity with quality.
“I don't think we're necessarily passing judgment on that here in this work,” Druck said.
The divergence result likewise has a specific meaning: Opus 5.5’s aggregate word usage comes closer to the human web articles Graphite selected. It does not establish whether readers would mistake an individual article for human work, or whether the model produces more accurate, original or useful content.
Old habits fade as others grow
Even as some familiar giveaways diminish, others become more common.
Compared with Opus 5, the new model uses “can help you” eight times as often and “is especially helpful” 12 times as often in Graphite’s articles. Opus 5 uses “is genuinely” 26 times as often as its successor.
The comparison between model families reveals another distinction. Opus 5.5 uses “the most powerful” 24 times as often as Astra, while Astra uses “not necessarily” 17 times as often as Opus 5.5. The researchers describe greater use of superlatives in Claude and more qualification of claims in Astra.
The total number of identified tells declines only modestly between the two latest Opus versions: 2,548 for Opus 5.5, compared with 2,666 for Opus 5. Across the broader Claude sequence, Graphite counts 3,746 in Opus 4 and 3,059 in Opus 4.6.
Those counts include words, short phrases and recurring patterns containing variable gaps. To qualify, a feature must appear at least twice the human rate after adjusting for text volume and must meet minimum frequency requirements. Its appearance in one document does not establish that AI wrote it.
Graphite separately tracks 11 familiar writing features, including formulaic endings and words commonly associated with AI. Their combined strength in distinguishing Opus 5.5 from the human sample falls 6% relative to Opus 5 and 53% relative to Opus 4.
The two results can coexist: a model can become more similar to human writing overall while retaining thousands of disproportionately common expressions.
Human writers are adapting, too
Druck said the public fixation on AI writing habits may also influence people who write without the tools.
He described a Graphite writer joking about deliberately introducing imperfections to avoid being mistaken for AI. He also recalled a presenter pausing to explain that an em-dash on a slide came from a person.
“Anecdotally, people are saying that they're changing their writing habits in order to not sound like AI,” Druck said.
Graphite has not established that behavior through a systematic study. Druck said the team wants to investigate it further.
He also suspects AI developers deliberately target some widely recognized tells, although the analysis does not establish which training decisions cause the observed changes.
“My guess is they have sort of targeted evals and they're trying to remove certain things,” he said.
If both people and models adjust their habits, the expressions readers treat as proof of AI authorship may quickly lose their usefulness. Druck said public perceptions could lag behind model releases.
For companies producing marketing copy, support material or other business content, that suggests a practical evaluation problem: instructions tailored to one model’s weaknesses may need revisiting when the model changes.
A baseline test, with important limits
Graphite’s original methodology starts with 10,000 web articles published before ChatGPT launched in November 2022. The researchers use GPT-4.1 to summarize each article, then ask other models to generate articles from those summaries at approximately the original length.
The approach keeps topics aligned without asking the models to reproduce the originals directly. Adding Opus 5.5 reduces the complete shared set to 9,974 topics, which means some figures for earlier models differ slightly from the original study.
All models receive one fixed prompt. The team does not specifically instruct them to imitate human authors or suppress known tells.
Druck said Graphite wants to follow up by testing instructions intended to make generated text sound more human, then examining which differences remain. That work could help distinguish easily changed vocabulary habits from more persistent features.
One candidate is variation in sentence length. Druck pointed to the original analysis showing that human articles vary their sentence lengths more than model-generated articles.
“Humans write some short sentences and some really long sentences,” he said. “The LLMs tend to generate sentences that are just more uniform in length.”
Whether that difference survives explicit instructions to humanize the output remains a hypothesis, rather than a finding of the current update.
Other limits also matter. A different prompt, human editing or a different task could change the results. The human comparison articles are older than the generated ones, so changes in web writing since 2022 could contribute to differences attributed to AI. And the model-based assessment of ornate prose may introduce evaluator bias.
Data others can examine
Smith emphasized that Graphite makes the articles and measured tells available for download, allowing others to examine the comparisons and pursue additional questions.
“It is something that anyone else can also look at and dig into and find other patterns as well,” he said.
The researchers want to continue tracking releases, but have not committed to testing every model. Druck suggested updates every few months might prove more manageable; Smith expressed interest in covering major releases.
For enterprise teams, the current study supports evaluating writing against the actual purpose of the content: whether it communicates clearly, makes defensible claims and fits the intended voice.
A reduction in conspicuous AI habits can help, but Graphite’s results show how much a model’s other preferences can persist—or change—between versions.
