How to Fact-Check a Coworker's AI Work (and Raise It Well)
A coworker's ChatGPT draft has errors. How to verify AI output fast, document exactly what's wrong, and raise it without becoming the office villain.
See Claude set up for your job
Skip the theory โ pick your profession and get the real workflows, ready-to-use prompts, and exact setup for your work.
TL;DR. Do not argue about whether it was written by AI. Check the claims against the source, write down exactly which ones do not match, and raise it as a document problem, not a person problem. The escalation script is below. Whoever sent the work owns it, so the goal of the conversation is to get the file corrected before a client acts on it.
You fact-check a coworker's AI work the same way you would check any draft: list the claims, trace them to the source, and write down what does not match. You raise it by taking the finding to the coworker first, in private, as a problem with the document rather than with them. And you protect yourself by keeping a dated note of what you found and who you told, because if the error turns out to be one of many, the process you followed is what everyone will look at.
That last point is not hypothetical. In August 2026 a post on r/careerguidance described how a single fact-check in a team meeting turned into a compliance review of a dozen client files. The details are in the fourth section. The lesson is not "stay quiet." It is that the check was right, the timing was unlucky, and a little documentation would have made it a non-event.
How do you tell if a coworker's work was written by AI?
You mostly do not need to, and you should not try to prove it. The question that matters is not "did they use ChatGPT" but "are the specifics in this document right." A human can be wrong about a coverage limit too. What AI changes is the pattern: the errors arrive wrapped in confident, well-structured prose, so they pass a skim that would have caught a sloppy human draft.
That said, some things do raise the odds that a document was generated without being read:
- It answers a question about a specific file without citing anything from the file. No policy number, no clause reference, no date. Just fluent generalities about what such a document "typically" contains.
- It is longer and more polished than the person's normal writing, with a summary paragraph that repeats the body.
- It came back too fast for a question that would have required opening a file.
- The sender cannot answer a follow-up. If one specific question gets "let me check and get back to you," they did not write it, or did not read it.
The Stanford and BetterUp researchers who coined the term workslop for this kind of output found that it shifts the burden downstream to the receiver, who has to redo the work or decide whether to trust it. Recognizing the pattern is useful. But the surface tells of AI text are a separate topic, covered in Why your AI writing sounds AI-written. For this guide, skip the forensics and go straight to the claims.
How do you fact-check AI output quickly?
- List every checkable claim in the document.
- Trace each claim to a primary source: the file, the contract, the policy, the original data.
- Check hardest where AI fails most: numbers, dates, names, exclusions and exceptions, citations.
- Mark anything you cannot trace as unverified, not wrong.
That is the whole method. It fits on a sticky note, and for a two-page email it takes about ten minutes. A few notes on making it fast.
Extract the claims with AI, verify them without it. The paste-able prompt in the FAQ below turns a draft into a numbered table of claims in seconds. Use it. Then close the chat window, because the model cannot open your client's file, and the whole point is to compare the claim against the file.
Exclusions and exceptions deserve their own line. A model asked about a policy, contract, or procedure will tend to describe the standard version of that document. The thing that makes yours different, the carve-out or the endorsement or the negotiated clause, is exactly what it will get wrong. If the draft says "this is excluded" or "this does not apply," that claim goes to the top of the list.
Unverified is not the same as wrong. A number you cannot find in the file may still be right. Your finding is "I could not source this," which is honest, defensible, and much easier to raise with a colleague than "you made this up."
Write down what you found as you go. Claim, what the source actually says, where in the source. This is your evidence for the next section, and it turns a fact-check from an opinion into a document.
For longer documents, the human review rubric for AI documentation is the scored version of this pass. If the draft is your own, the six-minute review is the right tool, and this guide will not repeat it. This is about someone else's output, after the fact, and the hard part is not the checking. It is the conversation.
How do you raise an AI mistake with a colleague or your manager?
Alison Green's advice column has been fielding the "my coworker uses AI badly" letter regularly since early 2026, and her answer is consistent: raise it as a work-quality problem with the person, then escalate on the substance if it is not fixed. The AI part is context, not the accusation. Here is a script built on that principle.
To the colleague, privately, before anyone else.
"I was going over the reply to [client] and a couple of the specifics did not match the file. The exclusion on page 2 of the draft says X, but the policy document says Y, and the limit quoted is [number] where the file says [number]. I wanted to flag it to you first so it can get corrected before anything goes out. I have the comparisons if it is useful."
Notice what is not in it: no "did you use ChatGPT," no "I figured this was AI." Lead with the discrepancy and the source. Most people, offered a way to fix an error quietly, will take it. If they say they used AI, that is your opening to suggest the four-step check, not to lecture.
To your manager, only if it is not corrected, or if the client has already acted on it.
"I want to flag a document accuracy issue. In [document] going to [client], two specifics did not match the source file, and I have noted the comparisons. I raised it with [colleague] on [date]. I am raising it with you because [it has not been corrected / the client may already have relied on it / I think the same issue may exist in other files]. What would you like me to do with the comparisons?"
Facts, dates, and a question. You are not asking for a verdict on your colleague. You are asking what to do with a finding, which is the manager's job to decide.
What to write down, the same day.
- The document, the date, and who sent it.
- Each claim that did not match, what the source says, and where.
- Who you told, when, and in what words. One line is enough.
- What happened next.
Keep it in your own notes, dated, in plain language. It is not a case file against anyone. It is the record that shows you found a problem, raised it in the right order, and did not sit on it.
Three things not to do. Do not fix it silently, because the wrong version is still in the file. Do not raise it in a group setting if you can avoid it, because that is where fact-checking turns into public correction. And do not present an AI-detector score, because it proves nothing and shifts the conversation from "is this right" to "did you cheat."
Who is responsible when AI makes a mistake at work?
The person who sent it. Every major vendor's terms put verification on the user, employers hold the employee who signed the email, and nobody has yet successfully argued "the model said so" as a defense. We covered the liability side in detail in Who is responsible when an AI makes a costly mistake. The short version for the workplace: a document with your name on it is yours, however it was drafted.
Which is why the August 2026 r/careerguidance thread, titled "Accidentally got half my team put under review because I fact-checked a coworker's AI use, what do I do?", turned out the way it did.
The poster worked at an insurance brokerage and had tested ChatGPT for client emails. It kept giving, in their words, "confident-sounding but wrong numbers when clients asked about coverage limits," so they stopped using it for that. Then, in a team meeting, a senior coworker mentioned she had used ChatGPT to answer a client's question about a specific policy exclusion because she "didn't feel like digging through the file." The poster said, in front of the team, that the exclusion she described was not how their policies worked. The team lead checked and confirmed it.
That should have been the end of it. Instead the director pulled the thread and "found the same issue in at least a dozen other client files going back months": wrong exclusions, wrong limits, AI-generated, never checked. Three senior coworkers were required to redo compliance training. The poster was asked to help write a quick-reference document on policy basics, and describes a "well look who's the favorite now" vibe from a couple of people who now sit next to them every day.
The top comment on the thread, from u/ironicmirror with 3,988 votes, put the underlying problem in one line: "The problem with AI is it makes dumb people look smart and make smart people feel dumb."
We are not going to tell an insurance brokerage how to run its compliance program, and neither should any blog. What the story shows is about process, and it generalizes to any field:
- The fact-check was correct and necessary. A dozen client files were wrong. The alternative to speaking up was not peace, it was a client discovering the error later.
- The timing turned a correction into an exposure. Said in a meeting, it became public. Said in a message to the coworker beforehand, it would have been a two-line fix, and the director would probably still have found the wider problem, just without a villain.
- The responsibility landed where it always lands. Not on the vendor, not on the tool, and not on the person who caught it. On the people who sent unchecked work to clients, and on a team with no review step for AI output.
- The absence of a rule was the real failure. Three people were retrained because nobody had written down "AI-drafted answers about a specific file get checked against the file." That sentence is the cheapest policy your team will ever adopt.
If you are the poster, the guilt is misplaced. Keep doing what you did, one step earlier and one step quieter. If you are the manager, write that sentence down before your own dozen files turn up.
Where this fits
Verification is the first of the two skills that separate being good at AI from just using it, and it is the one this guide is really about. The six-minute review covers your own drafts before they go out. This guide covers the case nobody planned for: the draft is already out, it is not yours, and you are the one who noticed.
Sources
- "Accidentally got half my team put under review because I fact-checked a coworker's AI use, what do I do?", r/careerguidance thread by u/Kind-King1491, August 2026, as reproduced by FAIL Blog (Cheezburger). The original Reddit permalink was not retrievable at the time of writing. A second reproduction is at JobAdvisor.
- Alison Green, "My coworker is using AI to do her work (badly)", Ask a Manager, January 2026.
- Kate Niederhoffer, Gabriella Rosen Kellerman, et al., "AI-Generated 'Workslop' Is Destroying Productivity", Harvard Business Review, September 2025.
- OpenAI, "Why language models hallucinate", September 2025.
- OpenAI Help Center, "Does ChatGPT tell the truth?"
See Claude set up for your job
Skip the theory โ pick your profession and get the real workflows, ready-to-use prompts, and exact setup for your work.
Set up AI for your job โ free, in about 2 minutes
Pick your profession and get your first working AI tool, a step-by-step guide, and a $0 plugin to take home. No credit card.
Get my free setupSee Claude set up for your job
Real workflows and ready-to-use prompts, profession by profession.
Frequently asked questions
Does AI make mistakes, and how often?+
Yes. Current models still fabricate specifics (figures, citations, policy terms, dates) with full confidence, and the rate is low per claim but not zero. In a client-facing document with thirty checkable claims, a low per-claim error rate still means a wrong claim is likely somewhere, so any document that goes to a client, a regulator, or a file of record needs a claim-by-claim check. The error rate also rises sharply when the model is asked about something specific to your organization that it has never seen, such as the terms of one particular contract or policy.
Can I trust ChatGPT answers for client work?+
Trust it for drafting, structure, and summarizing material you paste in. Do not trust it as a source of facts about your client, your policies, or your contracts, because it has never read them unless you gave them to it. If a coworker used ChatGPT to answer a client question about a specific document without opening the document, the answer is a guess dressed as a fact, and it needs to be checked against the file before anyone relies on it.
What are common AI hallucination examples in professional documents?+
The recurring ones are: a plausible-looking statute, clause, or section number that does not exist; a statistic with no traceable source; a coverage limit, deductible, fee, or deadline stated as fact when the model never saw the document; a case name or study that cannot be found; a quotation nobody said; arithmetic that is confidently wrong; and a summary of an attachment that describes what such a document usually says rather than what this one says. Each of them reads as more authoritative than a vague statement, which is exactly why they slip through.
Is there a ChatGPT fact-check prompt I can paste?+
Yes. Paste the draft and ask: 'List every factual claim in this text as a numbered table with three columns: the claim, whether it can be verified from the source material I have pasted, and what primary source would be needed to verify it if not. Do not tell me whether a claim is true. Do not rewrite the text. Flag any number, date, name, citation, exclusion, or exception.' The model is good at extracting the claims. It is not a reliable judge of whether they are true, so you still trace each flagged claim to the actual source yourself.
Who is responsible for AI mistakes: the employee, the manager, or the vendor?+
In practice, the person who sent the work owns the error, the same as if they had typed it themselves. The vendor's terms of use put verification on the user, and no employer accepts 'the chatbot said so' as a defense. Managers share responsibility when they knew AI was in use and set no review expectation, which is why the story in this guide ended with retraining for the whole group rather than one person. The practical rule: whoever's name goes on the document is responsible for every claim in it.
Should I use an AI fact checker tool?+
Use one to flag, not to verify. AI fact-checker tools are good at pulling out claims and pointing at likely problems, and they are useful as a first pass on a long document. They cannot open your client's file, your contract, or your internal policy, so they cannot tell you whether a number specific to your work is right. The check that actually settles it is the manual one: list every checkable claim, trace each to a primary source, look hardest at numbers, dates, names, exclusions, and citations, and mark anything you cannot trace as unverified rather than wrong.
Related Guides
The $10 AI Stack: Can Kimi K3 + DeepSeek Beat a $200 Plan?
The viral '$10 a month' Kimi K3 + DeepSeek stack, priced and limit-checked: what it really costs, who can actually buy it, what it does well for analysts, and when a $200 plan is still worth it.
Grok 4.6 for Professionals: What It's Good At, What It Isn't, and the Data Question (September 2026)
Grok 4.6 is xAI's (now SpaceXAI's) current flagship, released August 12, 2026. It is a serious, cheap coding and agent model, a weak writing model, and a product with data defaults that compliance teams do not like. Here is the straight answer on what it does well, what it costs, and whether client information belongs in it.
What Is Grok Bot, and Should It Sign Into Your Work Apps?
Grok Bot's AI teammates run on their own cloud computer and sign into your apps. What they do, what they cost, and what to check before connecting data.