A Third of New Web Pages Show AI Authorship, Pew Finds — What Professionals Need to Know About Credibility
Pew Research Center analyzed 490,000 web pages and found more than a third of content published since ChatGPT's launch shows signs of AI authorship. Commercial sites run ten times the AI-text rate of .edu or .gov pages. Here's what the shift means for professional credibility, disclosure, and standing out.
See Claude set up for your job
Skip the theory — pick your profession and get the real workflows, ready-to-use prompts, and exact setup for your work.
Models change every month.
One short weekly update that keeps this call current — free.
The largest web-content study of its kind just landed: Pew Research Center analyzed approximately 490,000 English-language web pages and found that more than a third of all pages published since ChatGPT's launch in November 2022 show signs of AI authorship. In a snapshot of July 2026 pages specifically, around one in ten carried detectable AI-text patterns. For professionals whose credibility depends on a recognizable human voice, this is the new baseline.
What Pew actually measured
The study used Open Pangram, a machine-learning model that identifies statistical patterns in word choice and sentence structure — the subtle fingerprint large language models leave in text at scale. Across nearly half a million pages collected by the Common Crawl archive between January 2021 and July 2026:
- Over one-third of post-ChatGPT pages show signs of AI authorship
- .com domains: ~10% of pages in July 2026
- .org domains: 4.6%
- .edu and .gov: ~1% each
Commercial websites are roughly double the rate of .org sites and ten times the rate of academic or government pages.
The researchers were careful to note limitations: AI detection models can misclassify human-written content, and the patterns they track can appear in human writing. The study's power is in tracking directional trends across hundreds of thousands of pages — not in providing reliable per-document verdicts.
The linguistic tells AI text leaves behind
Pew didn't just count AI-flagged pages. They tracked the specific features that have surged since 2022:
- Em dashes mid-sentence — like this one — roughly doubled in frequency
- Oxford commas (the comma before "and" in a series) rose 63%
- Words like delve, pivotal, crucial, and moreover more than doubled
- The construction "it's not just X, it's Y" nearly tripled
These patterns emerge because today's major models — trained on massive text datasets — developed consistent stylistic habits that bleed through even when you write with AI rather than asking it to write for you. Running a draft through ChatGPT to tighten the prose can leave the same statistical fingerprint as having ChatGPT write the draft from scratch.
What this actually changes for professionals
The credibility risk isn't detection — it's homogeneity. When a third of what your professional audience reads shows the same stylistic fingerprints, they start pattern-matching those patterns as "unchecked AI." A report that opens with "delve into the multifaceted landscape" now reads as generative output even if every word came from a human. The fix isn't to hunt and remove em dashes — it's to anchor your writing in what AI can't generate at scale: your specific data, your named client outcomes, your first-hand observations.
Trust signals are shifting toward specificity. The content that builds professional credibility in 2026 is content only you could have written: "We ran this approach on Q2 client renewals and saw a 23% lift" instead of "AI tools can significantly enhance operational efficiency." Specific, verifiable, first-person claims are both the credibility signal and the detection-immune signal.
Disclosure norms are evolving, but aren't settled. Major professional associations in law, communications, and HR have issued guidance in 2026, but standards vary by field, employer, and document type. The emerging professional default: disclose when AI shaped the substance — not when it only handled grammar or formatting. "Drafted with AI assistance" carries professional weight. "Checked with an AI grammar tool" generally doesn't need to be disclosed at all.
The bigger picture: bots reading AI-written content
Last month, Cloudflare confirmed that bots and AI agents now generate 57.4% of all web traffic — more requests come from automated systems than from human readers. The Pew data and the Cloudflare data together describe a loop already running: AI agents crawl web pages that are increasingly AI-written, summarize them for human readers, who receive a distilled version of a distilled version.
For professionals who publish content — whether that's a LinkedIn article, a client newsletter, or a firm blog — the practical implication is that the original you publish increasingly competes not just for human attention but for AI-agent summarization. The content that survives that process most reliably is specific, verifiable, and grounded in named human sources. Not "human-sounding" in a stylistic sense, but substantively human-sourced in a way models can't reproduce.
What to do this week
1. Audit your recent published work. Read your last two or three pieces aloud. Note every em dash inside a clause, every delve, every "it's not X, it's Y" formulation. These aren't errors — but you should know how your output reads in 2026's context.
2. Anchor your next piece to first-hand specifics. Include at least two claims that only you could have made: a project outcome, a client scenario (anonymized appropriately), a number you observed directly. That specificity is your credibility signal.
3. Know your field's disclosure standard. If your professional association or employer hasn't issued AI-content guidance, check the relevant major bodies in your field. The American Bar Association, the Society for Human Resource Management, and the International Association of Business Communicators have all issued guidance in 2026. Even if your employer hasn't, their expectations are likely shaped by one of these.
Sources
- How Much of the Internet Is Written With AI? — Pew Research Center, August 20, 2026
- A Third of Web Pages Published Since ChatGPT's Launch Show Signs of AI Authorship, Study Finds — TechCrunch, August 20, 2026
- A Third of the Post-ChatGPT Web Is AI-Written, Pew Finds — Decrypt, August 2026
See Claude set up for your job
Skip the theory — pick your profession and get the real workflows, ready-to-use prompts, and exact setup for your work.
Set up AI for your job — free, in about 2 minutes
Pick your profession and get your first working AI tool, a step-by-step guide, and a $0 plugin to take home. No credit card.
Get my free setupSee Claude set up for your job
Real workflows and ready-to-use prompts, profession by profession.
Frequently asked questions
What exactly did the Pew study find about AI-written web content?+
Pew Research Center analyzed approximately 490,000 English-language web pages from the Common Crawl archive, covering January 2021 through July 2026. They used Open Pangram, a machine-learning AI detection model, to identify statistical patterns in word choice and sentence structure. The result: more than a third of pages published after ChatGPT launched in November 2022 show signs of AI authorship. In a snapshot of pages from July 2026 specifically, about 10% carried these markers — with .com domains running roughly double the .org rate and about ten times the .edu or .gov rate.
Can AI detection tools accurately tell if text was written by AI?+
Not reliably on any individual document. Pew was explicit about this: Open Pangram, the model they used, can misclassify human-written pages as AI-generated. The study's value is in tracking statistical trends across hundreds of thousands of pages over time — not in offering a reliable per-document test. The same linguistic patterns that flag AI text (em dashes, 'delve,' 'it's not X, it's Y') can appear in human writing. This means you can't prove you wrote something by passing a detector test, and you shouldn't assume a positive detection result is accurate either.
Does AI-written content affect professional credibility?+
The credibility risk isn't from a detector catching you — it's from your writing sounding like everything else. When a third of what professionals read shows the same stylistic fingerprints, audiences start pattern-matching unconsciously. A report that leads with 'delve into the multifaceted landscape of...' now reads as unchecked AI output even if the author wrote every word. The content that sustains professional credibility — named case studies, your own data, specific project outcomes — is also what AI can't generate at scale. Those specifics are both the credibility signal and the detection-proof signal.
Should I disclose when I use AI to write professional content?+
The standard depends on your field, employer, and context, but the emerging professional norm in 2026 is to disclose when AI shaped the substance of a document — not when it only handled grammar or formatting. 'Drafted with AI assistance' means something meaningfully different from 'checked with an AI grammar tool.' Your clients, employers, and readers care about the former because it affects how much of the thinking and judgment is actually yours. Major professional associations in law, communications, and HR have issued guidance on this in 2026 — check the one most relevant to your field.
Why do .edu and .gov sites show such low AI-text rates?+
Pew found both at roughly 1%, versus ~10% for .com domains in July 2026. The most likely explanation is structural: academic publishing requires citation, peer review, and reproducibility — all of which create strong disincentives for generating text at scale. Government communications carry legal accountability and public-records obligations that push similarly toward human authorship. Commercial publishing has no equivalent constraint, which is why .com content shows the highest AI-text rate by far.
Related Guides
9 Workplace Monitoring Apps All Share Your Data. Here's Which Ones — and What Gets Sent.
A joint study from Columbia Law School, Northeastern, Vanderbilt, and UC Berkeley tested nine widely-used 'bossware' platforms — Hubstaff, Time Doctor 2, Deputy, and six others — and found all of them share worker names, emails, and employer info with Facebook, Google, Microsoft, and Yandex. Here's exactly what gets sent, where it goes, and what professionals can do today.
90% of Executives Say AI Hasn't Boosted Productivity — and AI Layoffs Are Making It Worse
A Federal Reserve survey of ~750 executives finds 90% say AI hasn't boosted productivity at their companies. A University of Pittsburgh study explains why: companies cutting jobs to offset AI costs are creating the employee resistance that kills the productivity gains they're chasing.
ChatGPT Can Now Read and Send Your iMessages on Mac — Should You Enable It?
OpenAI launched an Apple Messages plugin for ChatGPT Work and Codex on Mac on August 20, 2026. It can search your texts, draft replies, and send messages on your behalf. Here's what access it requires, what the real privacy tradeoffs are, and whether professionals should turn it on.