Five Percent — Research
Five Percent — Graphite
Research

AI Tells

AI Tells

Astra’s Em-Dash Rate Is 88% Below the Pre-ChatGPT Human Rate

Em Dash Use Across GPT Model Versions

Astra Uses Well-Known AI Tells 29% Less Often Than Human Writers

Well-Known AI Tell Use Across GPT Model Versions

Astra Talks About What Things Aren’t 12× as Often as Human Writers

Corrective Framing Across GPT Model Versions

Claude’s Mannered-Prose Score Has Risen 19% Since Opus 4

Mannered-Prose Scores Across Claude Model Versions

Claude Makes “Less X, More Y” Comparisons 82× as Often as Human Writers

“Less X, More Y” Comparisons Across Claude Model Versions

Gemini Uses Intensifiers 13× as Often as Human Writers

Intensifier Use Across Gemini Model Versions

Source: Graphite research, AI Tells (2026). The analysis uses 10,000 human articles and 90,000 AI-generated articles from nine models on the same topics. Values are usage rates or scores relative to human writing (1×). Mannered prose uses a 1,000-topic subsample. Composite measures include well-known tells (“delve,” “landscape”), corrective framing (“rather than,” “not just”), and intensifiers (“absolutely,” “highly”).

Key Takeaways

  • We identify AI tells at scale using 10,000 human articles and 90,000 AI-generated articles on the same topics (10,000 from each of nine models).

  • We find nearly 13,000 tells beyond the em dash.

    • Claude Opus 5, for example, uses the pattern “less like a _ and more like” about 105 times as often as human writers.
  • Model families have different tells. 65% of tells are unique to one model family.

    • GPT-6 Astra favors corrective phrasing like “rather than” and “not simply,” Claude Opus 5 favors superlative phrasing like “single most,” and Gemini 3.1 Pro favors formal transitions like “furthermore.”
    • Claude Opus 5 uses em dashes at a similar rate to human writers, while GPT-6 Astra and Gemini 3.1 Pro use them much less often.
    • Claude Opus 5 and Gemini 3.1 Pro have mannered-prose scores about 2.5 times the human score, compared with 1.2 times for GPT-6 Astra.
  • GPT-6 Astra follows the same pattern as earlier models. Tells shift between model versions, but they do not go away.

    • We found around 3,700 tells in Astra’s writing, slightly more than in GPT-5.6 Sol’s. Only 45% of the tells found across the two models are shared.
    • GPT-5.6 Sol uses “in addition” 14 times as often as Astra, while Astra uses “need not” 17 times as often as Sol.
    • Astra’s overall word distribution is slightly further from human writing than Sol’s.
  • Well-known tells are becoming less prominent, but AI writing is not consistently becoming more human.

    • Across model families, the average strength of 11 well-known tells falls by 21% to 50%.
    • Claude Opus 5’s overall word distribution is closer to human writing than Claude Opus 4’s, while GPT-6 Astra’s is further than GPT-4.1’s.
  • The tells data and raw article data are available to download.

Explore Terms and Frames Across Model Versions

Source: Graphite research, AI Tells (2026). The analysis uses 10,000 human articles and 90,000 AI-generated articles from nine models on the same topics. Values show each model's usage rate relative to human writing (1×). Terms appear in at least 250 articles and frames in at least 120. Terms and frames with no human occurrences are omitted because they do not have a finite human-relative ratio.

Introduction

AI-Generated Text Is Everywhere

In a separate study, we show that half of the articles being published online are AI-generated. Sites like Reddit and arXiv are grappling with how to reduce the amount of AI-generated content they receive.

AI is also increasingly used in the workplace. About half of US employees now use it at least occasionally, up from a fifth two years ago, and one in eight use it daily.

AI Has Tells — Including the Em Dash — in Its Writing

Can we distinguish AI-generated text from human writing? AI detectors use machine learning to classify text as AI-generated or human-written. Some of the patterns they identify are model-specific. These detectors are not perfect, but our previous evaluation shows they can be highly accurate in some settings.

Now that people create and encounter AI-generated content frequently, they have started to notice patterns in AI writing. We call them tells. One well-known example of an AI tell is the overuse of the em dash (“—” or “--”). AI models have tended to use the em dash more often than people do, so people now use the presence of an em dash as a primitive AI detector.

A more recent (and funny) example is the overuse of “goblins” in GPT-5.1 to 5.5. This post about the issue gives some insight into how tells can surface as unintended consequences of training.

This Report Looks at Tells in Depth

In this report, we develop easy-to-interpret methods to identify tells at scale. We compare tells across models and examine how they change between model versions.

Specifically, we use a topically aligned parallel corpus of articles to identify words, phrases, and other patterns that are disproportionately present in text generated by a particular model. For full details, see the Methodology section.

The tells data and raw article data are available to download.

AI Writing Has Many Tells

To identify tells, we look for words and two- or three-word phrases that appear more often in AI-generated text than in human-written text. For example, “is genuinely” appears 1,021 times in articles generated by Claude Opus 5 and four times in human articles. We also look for patterns of common words separated by gaps of up to three less common words, with an underscore representing the gap. For example, “less like a _ and more like” appears 125 times in Claude Opus 5 articles and once in human articles. We normalize these counts for the amount of text in each corpus before comparing them. Finally, we compute other text features, such as the standard deviation of sentence length.

Across all nine models, we found 12,877 unique words, phrases, and frames whose usage rate was at least twice the human rate and that met our frequency thresholds. Each model has between 2,355 and 3,746 of these tells. Together, GPT-6 Astra, Claude Opus 5, and Gemini 3.1 Pro account for 7,043 unique tells.

We observed several recurring themes in AI tells.

Evaluative Adjectives

Words that rate something without describing it.

GPT-6 Astra

TellTimes the Human Rate
dependable59×
practical26×
meaningful13×

Claude Opus 5

TellTimes the Human Rate
deliberate26×
measured4.2×
steady

Gemini 3.1 Pro

TellTimes the Human Rate
incredibly18×
profound25×
immense22×

Transitions

Transition words and phrases that connect ideas.

GPT-6 Astra

TellTimes the Human Rate
another dimension117×
together these95×

Claude Opus 5

TellTimes the Human Rate
what comes next28×
looking ahead the33×

Gemini 3.1 Pro

TellTimes the Human Rate
furthermore the43×
ultimately this78×
additionally the8.6×

Definition by Contrast

Saying what something is by saying what it is not.

GPT-6 Astra

TellTimes the Human Rate
not simply157×
rather than relying187×
the _ is not simply576×

Claude Opus 5

TellTimes the Human Rate
rather than merely160×
rather than simply43×
less like a _ and more like105×

Gemini 3.1 Pro

TellTimes the Human Rate
is not just18×
instead it is29×
is not just a _ it is153×

Avoiding Tradeoffs

AI often promises one benefit without giving up another.

GPT-6 Astra

TellTimes the Human Rate
without requiring70×
without losing22×
without sacrificing6.4×

Claude Opus 5

TellTimes the Human Rate
without losing6.1×
without sacrificing4.3×
without compromising2.6×

Gemini 3.1 Pro

TellTimes the Human Rate
without sacrificing7.8×
without losing5.4×
without compromising5.2×

Flagging Importance

Telling the reader that something matters.

GPT-6 Astra

TellTimes the Human Rate
distinction matters∞ (0 in human articles)
matters because357×

Claude Opus 5

TellTimes the Human Rate
matters because132×
matters more than93×

Gemini 3.1 Pro

TellTimes the Human Rate
absolutely essential32×
remarkably19×

AI Uses a Wider Vocabulary Within Articles

MTLD measures vocabulary variety within an article while accounting for its length. A higher score means the article maintains greater vocabulary variety as it gets longer. By this measure, AI articles use a wider vocabulary than human-written articles. This does not mean that AI has more varied writing styles across articles. AI models also use longer words, averaging 5.2 to 5.7 characters per word compared with 4.9 in human writing.

AI Uses a Wider Vocabulary Within Articles

Lexical Diversity by Model

Source: Graphite research, AI Tells (2026). The analysis uses 10,000 human articles and 90,000 AI-generated articles from nine models on the same topics. Mean MTLD across the 9,984 aligned topics. MTLD measures vocabulary variety within each article, not differences in writing style across articles. Higher values indicate greater vocabulary variety within an article.

AI Sentence Lengths Vary Less

Sentence lengths vary much less in AI writing than in human writing.

AI Sentence Lengths Vary Less

Sentence-Length Variation by Model

Source: Graphite research, AI Tells (2026). The analysis uses 10,000 human articles and 90,000 AI-generated articles from nine models on the same topics. Mean of each article's sentence-length standard deviation across the 9,984 aligned topics.

Tells by Model

The explorer below shows the most disproportionately used words, phrases, frames, and features, along with LLM-summarized themes. For the stylistic features, the larger the absolute Cohen's d, the more strongly the feature distinguishes AI from human writing.

Source: Graphite research, AI Tells (2026). The analysis uses 10,000 human articles and 90,000 AI-generated articles from nine models on the same topics. Ranked by the ratio of the model's length-normalized usage rate to the human rate and limited to terms appearing in at least 500 model articles.

Human Writing Is More Personable and Informal

Human writing has its own tells:

  • Speaks in the first person, with “we all know” and “we’re going to.”
  • Addresses the reader directly, with “you can’t” and “did you know.”
  • Uses casual qualifiers, with “pretty” and “sort of.”
  • Tells small anecdotes, with “he said” and “years ago.”
  • Draws on personal experience, with “my favorite,” “mom,” and “last year.”
  • Expresses personal enthusiasm, with “love to” and “excited to.”

These tells cluster into several broader themes.

Human Writing Tell Themes

Source: Graphite research, AI Tells (2026). The analysis uses 10,000 human articles and 90,000 AI-generated articles from nine models on the same topics. An LLM grouped the measured tells into themes and wrote the descriptions.

The most discriminative features also show this informality. Human writers use exclamation marks more than a hundred times as often as each of the three current models, and they interrupt themselves with parenthetical asides far more often.

Human Writers Use More Exclamation Marks and Parenthetical Asides

Exclamation Mark and Parenthetical Aside Use by Model

Exclamation marks

Parenthetical asides

Source: Graphite research, AI Tells (2026). The analysis uses 10,000 human articles and 90,000 AI-generated articles from nine models on the same topics. Mean occurrences per 1,000 words across the 9,984 aligned topics.

ChatGPT, Claude, and Gemini Have Different Tells

Model families have different tells. Across all nine models, 65% of tells are unique to one model family.

GPT-6 Astra

Compared with Claude Opus 5, GPT-6 Astra is more likely to hedge or qualify a claim.

TellTimes the Claude Rate
may provide36×
can provide23×
not necessarily18×
rather than focusing12×

Claude Opus 5

Compared with GPT-6 Astra, Claude Opus 5 is more likely to frame things as unusually important or extreme.

TellTimes the Astra Rate
single most∞ (0 in Astra articles)
arguably the most∞ (0 in Astra articles)
every single112×
enormously45×

Gemini 3.1 Pro

Compared with GPT-6 Astra, Gemini 3.1 Pro uses more intensifiers, formal transitions, and formulaic guide language.

TellTimes the Astra Rate
incredibly1,835×
absolutely110×
furthermore∞ (0 in Astra articles)
additionally1,139×
this comprehensive guide∞ (0 in Astra articles)

Em Dash Use Varies by Model

The em dash is the best-known tell, but AI models no longer overuse it consistently. GPT-6 Astra uses em dashes about one-eighth as often as human writers, Claude Opus 5 uses them at about the human rate, and Gemini 3.1 Pro has nearly stopped using them.

Em Dash Use Varies by Model

Em Dash Use by Model

Source: Graphite research, AI Tells (2026). The analysis uses 10,000 human articles and 90,000 AI-generated articles from nine models on the same topics. Mean em dashes per 1,000 words across the 9,984 aligned topics.

More Differences Between Models

Below, we show the most disproportionately used words, phrases, frames, and features for each pair of current models, along with LLM-summarized themes.

Claude Opus 5 tells

GPT-6 Astra tells

Source: Graphite research, AI Tells (2026). The analysis uses 10,000 human articles and 90,000 AI-generated articles from nine models on the same topics. Each side is ranked independently by its length-normalized usage rate relative to the other side and limited to terms appearing in at least 500 articles.

Tells Appear and Disappear Between Model Versions

Consecutive versions within the same model family can have substantially different tells. Between the two latest versions of each model family, 55% to 72% of tells are unique to one version. A tell associated with one model version may disappear in the next, while new tells often appear.

The model versions shown below are ordered separately for each model family and were not necessarily released at the same time.

GPT

GPT-6 Astra follows the same pattern as earlier model releases. We found 3,687 tells in Astra’s writing, slightly more than in GPT-5.6 Sol’s, but only 45% of the tells found across the two models are shared. Sol uses “in addition” 14 times as often as Astra, while Astra uses “need not” 17 times as often as Sol. The tells shift but do not go away.

Across GPT versions, marketing language like “streamline” and “unlock” becomes less common, while hedging with “may” and corrective framing become more common overall. Hedging peaks in GPT-5.6 Sol before declining in GPT-6 Astra, while corrective framing continues to rise.

Astra Uses Marketing Language 73% Less Often Than GPT-4.1

Marketing-Language Use Across GPT Model Versions

Astra Uses “may” 3× as Often as GPT-4.1

“may” Use Across GPT Model Versions

Astra Talks About What Things Aren’t 12× as Often as Human Writers

Corrective Framing Across GPT Model Versions

Source: Graphite research, AI Tells (2026). The analysis uses 10,000 human articles and 90,000 AI-generated articles from nine models on the same topics. Values show length-normalized usage relative to human writing (1×). Marketing language includes “payoff,” “unlock,” “elevate,” “amplify,” “streamline,” “boost,” “leverage,” “wins,” and “tangible.” Corrective framing includes “rather than,” “not just,” “not simply,” and “does not.”

Claude

Across Claude versions, hype words become less common, while “less X, more Y” comparisons and concrete numbers and timeframes become more common.

Claude Is Exaggerating Less

Hype-Word Use Across Claude Model Versions

Claude Is Making More “Less X, More Y” Comparisons

“less like a _ and more like” Use Across Claude Model Versions

Claude Is Using More Concrete Numbers and Timeframes

Concrete Number and Timeframe Use Across Claude Model Versions

Source: Graphite research, AI Tells (2026). The analysis uses 10,000 human articles and 90,000 AI-generated articles from nine models on the same topics. Values show length-normalized usage relative to human writing (1×). Hype words include “groundbreaking,” “paramount,” “unprecedented,” “revolutionary,” “invaluable,” “profound,” and “exceptional.” Concrete numbers and timeframes include “twenty minutes,” “ten minutes,” “five minutes,” “a dozen,” “a handful,” “six months,” “a decade,” and “two or three.”

Gemini

Across Gemini versions, contractions nearly disappear, while intensifiers and precision language become more common.

Gemini Is Using Fewer Contractions

Contraction Use Across Gemini Model Versions

Gemini Is Using More Intensifiers

Intensifier Use Across Gemini Model Versions

Gemini Is Using More Precision Language

Precision-Language Use Across Gemini Model Versions

Source: Graphite research, AI Tells (2026). The analysis uses 10,000 human articles and 90,000 AI-generated articles from nine models on the same topics. Values show length-normalized usage relative to human writing (1×). Contractions include “it's,” “don't,” “you're,” “isn't,” “doesn't,” “wasn't,” “you'll,” “we've,” and “let's.” Intensifiers include “absolutely,” “entirely,” “highly,” and “incredibly.” Precision language includes “exactly,” “precisely,” and “specifically.”

AI Writing Is Often More Mannered Than Human Writing

Anthropic calls writing that substitutes metaphor and flourish for direct statement mannered prose. It includes phrases like “a dial worth turning” where “a parameter worth varying” would be more direct. Claude Opus 5 and Gemini 3.1 Pro have mannered-prose scores about 2.5 times the human score, while GPT-6 Astra’s mannered-prose score is much closer to the human score.

AI Writing Is Often More Mannered Than Human Writing

Mannered-Prose Scores Across Model Versions

GPT

Claude

Gemini

Source: Graphite research, AI Tells (2026). The analysis uses 10,000 human articles and 90,000 AI-generated articles from nine models on the same topics. Values show the mean mannered-prose score relative to human writing (1×) across a subsample of 1,000 topics. Claude Opus 5 rates the amount of mannered prose in each complete article on a scale from 0 to 100 using Anthropic’s definition. The score is a model-judged prevalence index, not a mechanical sentence count.

Em Dash Use Varies by Model Version

Changes in em-dash use vary by model family. Em-dash use falls sharply from GPT-5 to GPT-5.6 Sol and remains low in Astra. It nearly disappears between the two Gemini versions. Claude Opus 4.6 uses fewer em dashes than Opus 4, but Opus 5 brings them back to about the human rate. GPT and Gemini may have overcorrected for this well-known tell. Their latest versions use em dashes far less often than human writers.

Em Dash Use Varies by Model Version

Em Dash Use Across Model Versions

GPT

Claude

Gemini

Source: Graphite research, AI Tells (2026). The analysis uses 10,000 human articles and 90,000 AI-generated articles from nine models on the same topics. Values show each model's mean em dash use relative to human writing (1×).

More Changes Between Model Versions

Below, we show the words, phrases, frames, and features that differ most between consecutive versions in each model family, along with LLM-summarized themes.

GPT-6 Astra tells

GPT-5.6 Sol tells

Source: Graphite research, AI Tells (2026). The analysis uses 10,000 human articles and 90,000 AI-generated articles from nine models on the same topics. Each side is ranked independently by its length-normalized usage rate relative to the other side and limited to terms appearing in at least 500 articles.

Explore Tells

Below, you can search recurring words, phrases, and frames and compare their usage rates with human writing across model versions.

Explore Terms and Frames Across Model Versions

Source: Graphite research, AI Tells (2026). The analysis uses 10,000 human articles and 90,000 AI-generated articles from nine models on the same topics. Values show each model's usage rate relative to human writing (1×). Terms appear in at least 250 articles and frames in at least 120. Terms and frames with no human occurrences are omitted because they do not have a finite human-relative ratio.

You can also compare other text features across the same model versions.

Explore Any Feature Across Model Versions

Mean value per article

Source: Graphite research, AI Tells (2026). The analysis uses 10,000 human articles and 90,000 AI-generated articles from nine models on the same topics. Mean per article across the 9,984 aligned topics. Features are ordered by how strongly they distinguish human writing from any model.

Tells Are Not Necessarily Fading

As new models are released, we might expect their writing to become closer to human writing and harder to distinguish. However, our results are mixed. We compare AI and human writing in three ways: semantic similarity to paired human articles, the strength of well-known tells, and divergence in word and phrase usage.

Semantic Similarity to Human Articles Barely Changes

We compare the semantic similarity of each model's articles to the corresponding human article. Similarity changes slightly, with no consistent movement toward human-written articles across model versions.

Semantic Similarity to Human Articles Barely Changes

Semantic Similarity to Human Articles Across Model Versions

GPT

Claude

Gemini

Source: Graphite research, AI Tells (2026). The analysis uses 10,000 human articles and 90,000 AI-generated articles from nine models on the same topics. Mean cosine similarity between each AI article and its paired human article using OpenAI's text-embedding-3-small. Higher values indicate greater semantic similarity.

Well-Known Tells Are Becoming Less Prominent

We examine the following features that measure how often specific well-known tells appear:

  • Hallmark words like “crucial” and “delve”
  • Antithesis (“not X, it’s Y”)
  • Boilerplate phrases like “it is important to note” and “at its core”
  • Formulaic closes
  • Closing calls to action
  • Sycophantic phrases
  • Self-referential signposting
  • Hedging
  • Parenthetical asides
  • Two measures of list use

In all three model families, these tells are less prominent in the latest model than in the earliest model we tested. They are also becoming less frequent overall. Across the nine features measured per 1,000 words, the combined frequency falls by 41% to 86% from the earliest to the latest model in each family. This decline may reflect deliberate efforts by AI companies to remove these tells.

Well-Known Tells Are Becoming Less Prominent

Strength of Well-Known AI Tells Across Model Versions

GPT

Claude

Gemini

Source: Graphite research, AI Tells (2026). The analysis uses 10,000 human articles and 90,000 AI-generated articles from nine models on the same topics. Mean absolute Cohen's d against human writing across 11 predefined features. Larger values indicate that the features more strongly distinguish the model's writing from human writing. Across the nine features measured per 1,000 words, the combined frequency falls by 41% to 86% from the earliest to the latest model in each family.

The total number of tells falls by 29% for Claude and 32% for Gemini between the earliest and latest versions we test, while GPT’s rises by 48%. But fewer tells do not necessarily mean that a model’s overall word distribution is closer to human writing.

Only Claude’s Word Choice Is Becoming More Human

We compare probability distributions over words and phrases. Larger values mean greater differences in how the model and human writers use words and phrases. Claude’s divergence falls sharply and then rises slightly, while GPT’s and Gemini’s rise overall. Under this measure, only Claude’s writing becomes more difficult to distinguish from human writing.

Only Claude's Word Choice Is Becoming More Human

Word-Distribution Divergence From Human Writing Across Model Versions

GPT

Claude

Gemini

Source: Graphite research, AI Tells (2026). The analysis uses 10,000 human articles and 90,000 AI-generated articles from nine models on the same topics. Jensen-Shannon divergence between each model's writing and human writing, measured over a shared unigram vocabulary. Zero indicates identical usage and one indicates no overlap.

Every Model’s Writing Is Closer to Another Model Than to Human Writing

The divergences above compare each model with human writing. The matrix below also compares models with one another. Each cell shows the divergence in word or phrase usage between its row and column. Darker cells mean less similar usage. The matrix is symmetric and the diagonal is zero because each distribution is identical to itself.

Every Model's Writing Is Closer to Another Model Than to Human Writing

Word-Distribution Divergence Among Models and Human Writing

Source: Graphite research, AI Tells (2026). The analysis uses 10,000 human articles and 90,000 AI-generated articles from nine models on the same topics. Values show Jensen-Shannon divergence. Zero indicates identical usage and one indicates no overlap. Shading is relative to the largest value in each table.

Across unigrams, bigrams, and trigrams, every model's writing is closer to at least one other model than to human writing. The models do not form a single tight cluster or separate neatly by model family, but they share patterns that distinguish their writing from human writing. Because every model received the same prompt, some of that similarity may come from the task rather than the models themselves.

Methodology

Corpus

We build a topically aligned corpus of AI and human writing. We gather 10,000 articles from Common Crawl that were published before ChatGPT launched on November 30, 2022. To obtain AI articles on the same topic, we first generate a summary of the original article using GPT-4.1. Then we ask different LLMs to generate an article based on the summary. This “bottleneck” method, inspired by autoencoders, helps each AI-generated article cover the same subtopics as the original without simply copying it. We generate articles using GPT-4.1, GPT-5, GPT-5.6 Sol, GPT-6 Astra, Claude Opus 4, Claude Opus 4.6, Claude Opus 5, Gemini 2.5 Pro, and Gemini 3.1 Pro. We compare model versions only within the same series, so we do not include Claude Fable or Gemini 3.7 Flash. The n-gram and frame analyses use the 9,984 topics where all nine models and the human article are present. Feature and semantic similarity comparisons use the matched articles available for each comparison.

We use the system prompt:

We use the user prompt:

Tell Identification Algorithm

AI detectors most likely use neural network models to identify AI-generated content. However, neural networks are “black boxes” and hence difficult to interpret. Our goal is to use a simple, interpretable method to identify AI tells, not to build the most accurate AI detector.

Before running the analyses, we clean every article by removing bylines, cookie notices, and boilerplate calls to action, stripping inline artifacts such as URLs and dates, and normalizing quotation marks.

We extract n-grams (unigrams, bigrams, and trigrams) from the text. We also extract templates like is not _ it is, where the underscore represents one to three uncommon words. We call these frames. Specifically, we keep the 300 most common words across all corpora and replace each run of one to three other words with a blank. Frames can contain up to seven tokens and must begin and end with one of the common words. We keep a frame only when its usage ratio is at least three times the highest ratio of any individual word it contains.

To identify disproportionately used words, phrases, and frames, we measure p_ai(feature) / p_human(feature), the feature’s probability in AI-generated content divided by its probability in human-written content. Here, p is the count divided by the number of positions of that length in the corpus. For model-to-model comparisons, we use the same calculation with a model probability in the denominator. For the AI-tell counts reported in this article, we require a ratio of at least 2, meaning the feature is at least twice as common in AI-generated content. When the denominator is zero, we report the ratio as infinite. We require tells to appear in at least 500 distinct articles for unigrams, 250 for bigrams, and 100 for trigrams. We require frames to appear in at least 120 distinct articles.

To identify broader themes, we ask Claude Opus 4.7 to group the top tells for each comparison into five to eight stylistic themes. The model generates the theme names and descriptions and selects examples from the measured tells. We call these “LLM-summarized themes” to distinguish them from the themes and examples we select for discussion in the text.

We use two additional sets of features. The first set covers surface properties of the text: sentence length, lexical diversity (MTLD), readability scores, and how often each kind of punctuation appears. The second set targets patterns we expect from LLM writing: how often the model uses overused LLM words like “crucial” or “delve,” whether it includes bullet or numbered lists, whether it ends with a call to action or a formulaic close like “In conclusion…,” how much it hedges, whether it signposts itself, whether it repeats phrases, and whether it opens with flattery.

We rank the additional features by absolute Cohen's d, the difference between class means in units of pooled standard deviation. This lets us compare features measured on different scales, such as commas per 1,000 words, lexical diversity, and word length. A larger absolute value means a larger effect.

We measure mannered prose separately on a subsample of 1,000 topics. Using Anthropic’s definition, Claude Opus 5 rates the amount of mannered prose in an article on a scale from 0 to 100. The score is anchored to the approximate share of sentences containing mannered prose, but we treat it as a model-judged index rather than a mechanical count. We also evaluate about 2,000 articles with a more detailed phrase-extraction method. The two measures have a Spearman correlation of 0.75.

To measure semantic similarity, we embed each AI article and its paired human article using OpenAI’s text-embedding-3-small, calculate the cosine similarity for each pair, and report the mean for each model.

To compare how models and human writers use n-grams, we use Jensen-Shannon divergence. We calculate it from raw, unsmoothed counts over a shared vocabulary. It ranges from 0 to 1, where 0 means identical usage and 1 means no overlap. It is measured in bits. For unigram distributions, a value of 0.10 means that a randomly selected word provides an average of 0.10 bits of information about whether it came from model or human writing, treating each as equally likely.

Limitations

The tells we identify depend on how our corpus was constructed. Our documents are web articles. Tells may differ in other contexts, such as conversations with AI models.

We use articles published before the launch of ChatGPT because they are more likely to have been written by humans. AI articles were generated by newer models, so changes in writing style since 2022 may be misidentified as AI tells. Some human articles also contain CTAs and boilerplate that may remain after cleaning and affect the tells we identify.

We generate AI articles using a single, fixed prompt. A different prompt would likely change the writing and the tells we identify. For example, the human articles use contractions and the second person much more often. We could instruct the models to reproduce those traits, but that could hide some tells. Instead, we use a relatively general prompt without explicitly asking the models to imitate human web writing.

Claude Opus 5 assigns the mannered-prose scores without being told whether an article was written by a human or which model generated it. Because this is still a model-based judgment, some evaluator-specific bias is possible, although the task asks for a defined stylistic rating rather than a preference between outputs.

We also checked whether tells differ by topic and found similar tells across topic clusters. A larger corpus may be needed to identify topic-specific tells.

Our results are limited to the model versions we compared.

Identifying tells is fun, but we do not recommend classifying authorship based on individual tells. For this task, we recommend using an AI detector.

Conclusion

AI has many more tells than the em dash. Across all nine models, we found nearly 13,000 words, phrases, and frames used at least twice the human rate. Model families have different tells. 65% of tells are unique to one model family. Tells also shift between model versions.

GPT-6 Astra follows the same pattern as earlier models. Its tells differ from GPT-5.6 Sol’s, but they have not disappeared. Astra’s overall word distribution is also slightly further from human writing than Sol’s.

Well-known tells are becoming less prominent, but AI writing is not consistently becoming more human. Across the three model families, only Claude’s overall word distribution is becoming more similar to human writing.

Researchers

Gregory Druck, PhDChief AI Officer

Chief AI Officer at Graphite, leading the team building AI tools for growth and researching how AI is reshaping marketing. Previously Chief Data Scientist at Yummly and an NLP and search researcher at Yahoo! Research, with internships at Google and Microsoft. Ph.D. in machine learning from UMass Amherst, advised by Andrew McCallum.

Read Full Bio
Jose Luis Paredes, PhDSenior Data Specialist

Senior Data Scientist at Graphite. Ph.D. and Master's from the University of Delaware, followed by more than two decades as a professor at the University of Los Andes. Author of over 50 peer-reviewed research papers and holder of 5 U.S. patents. Previously Head of Data Science at GoToDigital.

Read Full Bio
Ethan SmithCEO

Founder and CEO of Graphite, the research-driven growth agency behind work for Webflow, Adobe, and Upwork. Teaches SEO and AEO at Reforge and is an adjunct professor at IE Business School. Research published in the Financial Times, Axios, and The Atlantic. Previously a growth advisor to Masterclass, Robinhood, and Honey.

Read Full Bio