AI Tells
AI Tells
Key Takeaways
We identify AI tells at scale using 10,000 human articles and 90,000 AI-generated articles on the same topics (10,000 from each of nine models).
We find nearly 13,000 tells beyond the em dash.
- Claude Opus 5, for example, uses the pattern “less like a _ and more like” about 105 times as often as human writers.
Model families have different tells. 65% of tells are unique to one model family.
- GPT-6 Astra favors corrective phrasing like “rather than” and “not simply,” Claude Opus 5 favors superlative phrasing like “single most,” and Gemini 3.1 Pro favors formal transitions like “furthermore.”
- Claude Opus 5 uses em dashes at a similar rate to human writers, while GPT-6 Astra and Gemini 3.1 Pro use them much less often.
- Claude Opus 5 and Gemini 3.1 Pro have mannered-prose scores about 2.5 times the human score, compared with 1.2 times for GPT-6 Astra.
GPT-6 Astra follows the same pattern as earlier models. Tells shift between model versions, but they do not go away.
- We found around 3,700 tells in Astra’s writing, slightly more than in GPT-5.6 Sol’s. Only 45% of the tells found across the two models are shared.
- GPT-5.6 Sol uses “in addition” 14 times as often as Astra, while Astra uses “need not” 17 times as often as Sol.
- Astra’s overall word distribution is slightly further from human writing than Sol’s.
Well-known tells are becoming less prominent, but AI writing is not consistently becoming more human.
- Across model families, the average strength of 11 well-known tells falls by 21% to 50%.
- Claude Opus 5’s overall word distribution is closer to human writing than Claude Opus 4’s, while GPT-6 Astra’s is further than GPT-4.1’s.
The tells data and raw article data are available to download.
Explore Terms and Frames Across Model Versions
Introduction
AI-Generated Text Is Everywhere
In a separate study, we show that half of the articles being published online are AI-generated. Sites like Reddit and arXiv are grappling with how to reduce the amount of AI-generated content they receive.
AI is also increasingly used in the workplace. About half of US employees now use it at least occasionally, up from a fifth two years ago, and one in eight use it daily.
AI Has Tells — Including the Em Dash — in Its Writing
Can we distinguish AI-generated text from human writing? AI detectors use machine learning to classify text as AI-generated or human-written. Some of the patterns they identify are model-specific. These detectors are not perfect, but our previous evaluation shows they can be highly accurate in some settings.
Now that people create and encounter AI-generated content frequently, they have started to notice patterns in AI writing. We call them tells. One well-known example of an AI tell is the overuse of the em dash (“—” or “--”). AI models have tended to use the em dash more often than people do, so people now use the presence of an em dash as a primitive AI detector.
A more recent (and funny) example is the overuse of “goblins” in GPT-5.1 to 5.5. This post about the issue gives some insight into how tells can surface as unintended consequences of training.
This Report Looks at Tells in Depth
In this report, we develop easy-to-interpret methods to identify tells at scale. We compare tells across models and examine how they change between model versions.
Specifically, we use a topically aligned parallel corpus of articles to identify words, phrases, and other patterns that are disproportionately present in text generated by a particular model. For full details, see the Methodology section.
The tells data and raw article data are available to download.
AI Writing Has Many Tells
To identify tells, we look for words and two- or three-word phrases that appear more often in AI-generated text than in human-written text. For example, “is genuinely” appears 1,021 times in articles generated by Claude Opus 5 and four times in human articles. We also look for patterns of common words separated by gaps of up to three less common words, with an underscore representing the gap. For example, “less like a _ and more like” appears 125 times in Claude Opus 5 articles and once in human articles. We normalize these counts for the amount of text in each corpus before comparing them. Finally, we compute other text features, such as the standard deviation of sentence length.
Across all nine models, we found 12,877 unique words, phrases, and frames whose usage rate was at least twice the human rate and that met our frequency thresholds. Each model has between 2,355 and 3,746 of these tells. Together, GPT-6 Astra, Claude Opus 5, and Gemini 3.1 Pro account for 7,043 unique tells.
We observed several recurring themes in AI tells.
Evaluative Adjectives
Words that rate something without describing it.
GPT-6 Astra
| Tell | Times the Human Rate |
|---|---|
| dependable | 59× |
| practical | 26× |
| meaningful | 13× |
Claude Opus 5
| Tell | Times the Human Rate |
|---|---|
| deliberate | 26× |
| measured | 4.2× |
| steady | 6× |
Gemini 3.1 Pro
| Tell | Times the Human Rate |
|---|---|
| incredibly | 18× |
| profound | 25× |
| immense | 22× |
Transitions
Transition words and phrases that connect ideas.
GPT-6 Astra
| Tell | Times the Human Rate |
|---|---|
| another dimension | 117× |
| together these | 95× |
Claude Opus 5
| Tell | Times the Human Rate |
|---|---|
| what comes next | 28× |
| looking ahead the | 33× |
Gemini 3.1 Pro
| Tell | Times the Human Rate |
|---|---|
| furthermore the | 43× |
| ultimately this | 78× |
| additionally the | 8.6× |
Definition by Contrast
Saying what something is by saying what it is not.
GPT-6 Astra
| Tell | Times the Human Rate |
|---|---|
| not simply | 157× |
| rather than relying | 187× |
| the _ is not simply | 576× |
Claude Opus 5
| Tell | Times the Human Rate |
|---|---|
| rather than merely | 160× |
| rather than simply | 43× |
| less like a _ and more like | 105× |
Gemini 3.1 Pro
| Tell | Times the Human Rate |
|---|---|
| is not just | 18× |
| instead it is | 29× |
| is not just a _ it is | 153× |
Avoiding Tradeoffs
AI often promises one benefit without giving up another.
GPT-6 Astra
| Tell | Times the Human Rate |
|---|---|
| without requiring | 70× |
| without losing | 22× |
| without sacrificing | 6.4× |
Claude Opus 5
| Tell | Times the Human Rate |
|---|---|
| without losing | 6.1× |
| without sacrificing | 4.3× |
| without compromising | 2.6× |
Gemini 3.1 Pro
| Tell | Times the Human Rate |
|---|---|
| without sacrificing | 7.8× |
| without losing | 5.4× |
| without compromising | 5.2× |
Flagging Importance
Telling the reader that something matters.
GPT-6 Astra
| Tell | Times the Human Rate |
|---|---|
| distinction matters | ∞ (0 in human articles) |
| matters because | 357× |
Claude Opus 5
| Tell | Times the Human Rate |
|---|---|
| matters because | 132× |
| matters more than | 93× |
Gemini 3.1 Pro
| Tell | Times the Human Rate |
|---|---|
| absolutely essential | 32× |
| remarkably | 19× |
AI Uses a Wider Vocabulary Within Articles
MTLD measures vocabulary variety within an article while accounting for its length. A higher score means the article maintains greater vocabulary variety as it gets longer. By this measure, AI articles use a wider vocabulary than human-written articles. This does not mean that AI has more varied writing styles across articles. AI models also use longer words, averaging 5.2 to 5.7 characters per word compared with 4.9 in human writing.
AI Uses a Wider Vocabulary Within Articles
Lexical Diversity by Model
AI Sentence Lengths Vary Less
Sentence lengths vary much less in AI writing than in human writing.
AI Sentence Lengths Vary Less
Sentence-Length Variation by Model
Tells by Model
The explorer below shows the most disproportionately used words, phrases, frames, and features, along with LLM-summarized themes. For the stylistic features, the larger the absolute Cohen's d, the more strongly the feature distinguishes AI from human writing.
Human Writing Is More Personable and Informal
Human writing has its own tells:
- Speaks in the first person, with “we all know” and “we’re going to.”
- Addresses the reader directly, with “you can’t” and “did you know.”
- Uses casual qualifiers, with “pretty” and “sort of.”
- Tells small anecdotes, with “he said” and “years ago.”
- Draws on personal experience, with “my favorite,” “mom,” and “last year.”
- Expresses personal enthusiasm, with “love to” and “excited to.”
These tells cluster into several broader themes.
Human Writing Tell Themes
The most discriminative features also show this informality. Human writers use exclamation marks more than a hundred times as often as each of the three current models, and they interrupt themselves with parenthetical asides far more often.
Human Writers Use More Exclamation Marks and Parenthetical Asides
Exclamation Mark and Parenthetical Aside Use by Model
Exclamation marks
Parenthetical asides
ChatGPT, Claude, and Gemini Have Different Tells
Model families have different tells. Across all nine models, 65% of tells are unique to one model family.
GPT-6 Astra
Compared with Claude Opus 5, GPT-6 Astra is more likely to hedge or qualify a claim.
| Tell | Times the Claude Rate |
|---|---|
| may provide | 36× |
| can provide | 23× |
| not necessarily | 18× |
| rather than focusing | 12× |
Claude Opus 5
Compared with GPT-6 Astra, Claude Opus 5 is more likely to frame things as unusually important or extreme.
| Tell | Times the Astra Rate |
|---|---|
| single most | ∞ (0 in Astra articles) |
| arguably the most | ∞ (0 in Astra articles) |
| every single | 112× |
| enormously | 45× |
Gemini 3.1 Pro
Compared with GPT-6 Astra, Gemini 3.1 Pro uses more intensifiers, formal transitions, and formulaic guide language.
| Tell | Times the Astra Rate |
|---|---|
| incredibly | 1,835× |
| absolutely | 110× |
| furthermore | ∞ (0 in Astra articles) |
| additionally | 1,139× |
| this comprehensive guide | ∞ (0 in Astra articles) |
Em Dash Use Varies by Model
The em dash is the best-known tell, but AI models no longer overuse it consistently. GPT-6 Astra uses em dashes about one-eighth as often as human writers, Claude Opus 5 uses them at about the human rate, and Gemini 3.1 Pro has nearly stopped using them.
Em Dash Use Varies by Model
Em Dash Use by Model
More Differences Between Models
Below, we show the most disproportionately used words, phrases, frames, and features for each pair of current models, along with LLM-summarized themes.
Claude Opus 5 tells
GPT-6 Astra tells
Tells Appear and Disappear Between Model Versions
Consecutive versions within the same model family can have substantially different tells. Between the two latest versions of each model family, 55% to 72% of tells are unique to one version. A tell associated with one model version may disappear in the next, while new tells often appear.
The model versions shown below are ordered separately for each model family and were not necessarily released at the same time.
GPT
GPT-6 Astra follows the same pattern as earlier model releases. We found 3,687 tells in Astra’s writing, slightly more than in GPT-5.6 Sol’s, but only 45% of the tells found across the two models are shared. Sol uses “in addition” 14 times as often as Astra, while Astra uses “need not” 17 times as often as Sol. The tells shift but do not go away.
Across GPT versions, marketing language like “streamline” and “unlock” becomes less common, while hedging with “may” and corrective framing become more common overall. Hedging peaks in GPT-5.6 Sol before declining in GPT-6 Astra, while corrective framing continues to rise.
Astra Uses Marketing Language 73% Less Often Than GPT-4.1
Marketing-Language Use Across GPT Model Versions
Astra Uses “may” 3× as Often as GPT-4.1
“may” Use Across GPT Model Versions
Astra Talks About What Things Aren’t 12× as Often as Human Writers
Corrective Framing Across GPT Model Versions
Claude
Across Claude versions, hype words become less common, while “less X, more Y” comparisons and concrete numbers and timeframes become more common.
Claude Is Exaggerating Less
Hype-Word Use Across Claude Model Versions
Claude Is Making More “Less X, More Y” Comparisons
“less like a _ and more like” Use Across Claude Model Versions
Claude Is Using More Concrete Numbers and Timeframes
Concrete Number and Timeframe Use Across Claude Model Versions
Gemini
Across Gemini versions, contractions nearly disappear, while intensifiers and precision language become more common.
Gemini Is Using Fewer Contractions
Contraction Use Across Gemini Model Versions
Gemini Is Using More Intensifiers
Intensifier Use Across Gemini Model Versions
Gemini Is Using More Precision Language
Precision-Language Use Across Gemini Model Versions
AI Writing Is Often More Mannered Than Human Writing
Anthropic calls writing that substitutes metaphor and flourish for direct statement mannered prose. It includes phrases like “a dial worth turning” where “a parameter worth varying” would be more direct. Claude Opus 5 and Gemini 3.1 Pro have mannered-prose scores about 2.5 times the human score, while GPT-6 Astra’s mannered-prose score is much closer to the human score.
AI Writing Is Often More Mannered Than Human Writing
Mannered-Prose Scores Across Model Versions
GPT
Claude
Gemini
Em Dash Use Varies by Model Version
Changes in em-dash use vary by model family. Em-dash use falls sharply from GPT-5 to GPT-5.6 Sol and remains low in Astra. It nearly disappears between the two Gemini versions. Claude Opus 4.6 uses fewer em dashes than Opus 4, but Opus 5 brings them back to about the human rate. GPT and Gemini may have overcorrected for this well-known tell. Their latest versions use em dashes far less often than human writers.
Em Dash Use Varies by Model Version
Em Dash Use Across Model Versions
GPT
Claude
Gemini
More Changes Between Model Versions
Below, we show the words, phrases, frames, and features that differ most between consecutive versions in each model family, along with LLM-summarized themes.
GPT-6 Astra tells
GPT-5.6 Sol tells
Explore Tells
Below, you can search recurring words, phrases, and frames and compare their usage rates with human writing across model versions.
Explore Terms and Frames Across Model Versions
You can also compare other text features across the same model versions.
Explore Any Feature Across Model Versions
Mean value per article
Tells Are Not Necessarily Fading
As new models are released, we might expect their writing to become closer to human writing and harder to distinguish. However, our results are mixed. We compare AI and human writing in three ways: semantic similarity to paired human articles, the strength of well-known tells, and divergence in word and phrase usage.
Semantic Similarity to Human Articles Barely Changes
We compare the semantic similarity of each model's articles to the corresponding human article. Similarity changes slightly, with no consistent movement toward human-written articles across model versions.
Semantic Similarity to Human Articles Barely Changes
Semantic Similarity to Human Articles Across Model Versions
GPT
Claude
Gemini
Well-Known Tells Are Becoming Less Prominent
We examine the following features that measure how often specific well-known tells appear:
- Hallmark words like “crucial” and “delve”
- Antithesis (“not X, it’s Y”)
- Boilerplate phrases like “it is important to note” and “at its core”
- Formulaic closes
- Closing calls to action
- Sycophantic phrases
- Self-referential signposting
- Hedging
- Parenthetical asides
- Two measures of list use
In all three model families, these tells are less prominent in the latest model than in the earliest model we tested. They are also becoming less frequent overall. Across the nine features measured per 1,000 words, the combined frequency falls by 41% to 86% from the earliest to the latest model in each family. This decline may reflect deliberate efforts by AI companies to remove these tells.
Well-Known Tells Are Becoming Less Prominent
Strength of Well-Known AI Tells Across Model Versions
GPT
Claude
Gemini
The total number of tells falls by 29% for Claude and 32% for Gemini between the earliest and latest versions we test, while GPT’s rises by 48%. But fewer tells do not necessarily mean that a model’s overall word distribution is closer to human writing.
Only Claude’s Word Choice Is Becoming More Human
We compare probability distributions over words and phrases. Larger values mean greater differences in how the model and human writers use words and phrases. Claude’s divergence falls sharply and then rises slightly, while GPT’s and Gemini’s rise overall. Under this measure, only Claude’s writing becomes more difficult to distinguish from human writing.
Only Claude's Word Choice Is Becoming More Human
Word-Distribution Divergence From Human Writing Across Model Versions
GPT
Claude
Gemini
Every Model’s Writing Is Closer to Another Model Than to Human Writing
The divergences above compare each model with human writing. The matrix below also compares models with one another. Each cell shows the divergence in word or phrase usage between its row and column. Darker cells mean less similar usage. The matrix is symmetric and the diagonal is zero because each distribution is identical to itself.
Every Model's Writing Is Closer to Another Model Than to Human Writing
Word-Distribution Divergence Among Models and Human Writing
Across unigrams, bigrams, and trigrams, every model's writing is closer to at least one other model than to human writing. The models do not form a single tight cluster or separate neatly by model family, but they share patterns that distinguish their writing from human writing. Because every model received the same prompt, some of that similarity may come from the task rather than the models themselves.
Methodology
Corpus
We build a topically aligned corpus of AI and human writing. We gather 10,000 articles from Common Crawl that were published before ChatGPT launched on November 30, 2022. To obtain AI articles on the same topic, we first generate a summary of the original article using GPT-4.1. Then we ask different LLMs to generate an article based on the summary. This “bottleneck” method, inspired by autoencoders, helps each AI-generated article cover the same subtopics as the original without simply copying it. We generate articles using GPT-4.1, GPT-5, GPT-5.6 Sol, GPT-6 Astra, Claude Opus 4, Claude Opus 4.6, Claude Opus 5, Gemini 2.5 Pro, and Gemini 3.1 Pro. We compare model versions only within the same series, so we do not include Claude Fable or Gemini 3.7 Flash. The n-gram and frame analyses use the 9,984 topics where all nine models and the human article are present. Feature and semantic similarity comparisons use the matched articles available for each comparison.
We use the system prompt:
We use the user prompt:
Tell Identification Algorithm
AI detectors most likely use neural network models to identify AI-generated content. However, neural networks are “black boxes” and hence difficult to interpret. Our goal is to use a simple, interpretable method to identify AI tells, not to build the most accurate AI detector.
Before running the analyses, we clean every article by removing bylines, cookie notices, and boilerplate calls to action, stripping inline artifacts such as URLs and dates, and normalizing quotation marks.
We extract n-grams (unigrams, bigrams, and trigrams) from the text. We also extract templates like is not _ it is, where the underscore represents one to three uncommon words. We call these frames. Specifically, we keep the 300 most common words across all corpora and replace each run of one to three other words with a blank. Frames can contain up to seven tokens and must begin and end with one of the common words. We keep a frame only when its usage ratio is at least three times the highest ratio of any individual word it contains.
To identify disproportionately used words, phrases, and frames, we measure p_ai(feature) / p_human(feature), the feature’s probability in AI-generated content divided by its probability in human-written content. Here, p is the count divided by the number of positions of that length in the corpus. For model-to-model comparisons, we use the same calculation with a model probability in the denominator. For the AI-tell counts reported in this article, we require a ratio of at least 2, meaning the feature is at least twice as common in AI-generated content. When the denominator is zero, we report the ratio as infinite. We require tells to appear in at least 500 distinct articles for unigrams, 250 for bigrams, and 100 for trigrams. We require frames to appear in at least 120 distinct articles.
To identify broader themes, we ask Claude Opus 4.7 to group the top tells for each comparison into five to eight stylistic themes. The model generates the theme names and descriptions and selects examples from the measured tells. We call these “LLM-summarized themes” to distinguish them from the themes and examples we select for discussion in the text.
We use two additional sets of features. The first set covers surface properties of the text: sentence length, lexical diversity (MTLD), readability scores, and how often each kind of punctuation appears. The second set targets patterns we expect from LLM writing: how often the model uses overused LLM words like “crucial” or “delve,” whether it includes bullet or numbered lists, whether it ends with a call to action or a formulaic close like “In conclusion…,” how much it hedges, whether it signposts itself, whether it repeats phrases, and whether it opens with flattery.
We rank the additional features by absolute Cohen's d, the difference between class means in units of pooled standard deviation. This lets us compare features measured on different scales, such as commas per 1,000 words, lexical diversity, and word length. A larger absolute value means a larger effect.
We measure mannered prose separately on a subsample of 1,000 topics. Using Anthropic’s definition, Claude Opus 5 rates the amount of mannered prose in an article on a scale from 0 to 100. The score is anchored to the approximate share of sentences containing mannered prose, but we treat it as a model-judged index rather than a mechanical count. We also evaluate about 2,000 articles with a more detailed phrase-extraction method. The two measures have a Spearman correlation of 0.75.
To measure semantic similarity, we embed each AI article and its paired human article using OpenAI’s text-embedding-3-small, calculate the cosine similarity for each pair, and report the mean for each model.
To compare how models and human writers use n-grams, we use Jensen-Shannon divergence. We calculate it from raw, unsmoothed counts over a shared vocabulary. It ranges from 0 to 1, where 0 means identical usage and 1 means no overlap. It is measured in bits. For unigram distributions, a value of 0.10 means that a randomly selected word provides an average of 0.10 bits of information about whether it came from model or human writing, treating each as equally likely.
Limitations
The tells we identify depend on how our corpus was constructed. Our documents are web articles. Tells may differ in other contexts, such as conversations with AI models.
We use articles published before the launch of ChatGPT because they are more likely to have been written by humans. AI articles were generated by newer models, so changes in writing style since 2022 may be misidentified as AI tells. Some human articles also contain CTAs and boilerplate that may remain after cleaning and affect the tells we identify.
We generate AI articles using a single, fixed prompt. A different prompt would likely change the writing and the tells we identify. For example, the human articles use contractions and the second person much more often. We could instruct the models to reproduce those traits, but that could hide some tells. Instead, we use a relatively general prompt without explicitly asking the models to imitate human web writing.
Claude Opus 5 assigns the mannered-prose scores without being told whether an article was written by a human or which model generated it. Because this is still a model-based judgment, some evaluator-specific bias is possible, although the task asks for a defined stylistic rating rather than a preference between outputs.
We also checked whether tells differ by topic and found similar tells across topic clusters. A larger corpus may be needed to identify topic-specific tells.
Our results are limited to the model versions we compared.
Identifying tells is fun, but we do not recommend classifying authorship based on individual tells. For this task, we recommend using an AI detector.
Conclusion
AI has many more tells than the em dash. Across all nine models, we found nearly 13,000 words, phrases, and frames used at least twice the human rate. Model families have different tells. 65% of tells are unique to one model family. Tells also shift between model versions.
GPT-6 Astra follows the same pattern as earlier models. Its tells differ from GPT-5.6 Sol’s, but they have not disappeared. Astra’s overall word distribution is also slightly further from human writing than Sol’s.
Well-known tells are becoming less prominent, but AI writing is not consistently becoming more human. Across the three model families, only Claude’s overall word distribution is becoming more similar to human writing.
Researchers
Chief AI Officer at Graphite, leading the team building AI tools for growth and researching how AI is reshaping marketing. Previously Chief Data Scientist at Yummly and an NLP and search researcher at Yahoo! Research, with internships at Google and Microsoft. Ph.D. in machine learning from UMass Amherst, advised by Andrew McCallum.
Read Full BioSenior Data Scientist at Graphite. Ph.D. and Master's from the University of Delaware, followed by more than two decades as a professor at the University of Los Andes. Author of over 50 peer-reviewed research papers and holder of 5 U.S. patents. Previously Head of Data Science at GoToDigital.
Read Full BioFounder and CEO of Graphite, the research-driven growth agency behind work for Webflow, Adobe, and Upwork. Teaches SEO and AEO at Reforge and is an adjunct professor at IE Business School. Research published in the Financial Times, Axios, and The Atlantic. Previously a growth advisor to Masterclass, Robinhood, and Honey.
Read Full Bio

