The evaluation pass
- Restate the briefwhat you asked, for whom, from which inputs
- Is it complete?every requirement and input covered
- Is it accurate?trace each claim to its source
- Weigh the stakeshow costly is an error here?
- Fix or ask againcorrect it yourself, or re-prompt precisely
Re-check the revised output the same way — a fix can introduce a new error
Two different questions
Accuracy asks whether what the output says is true: the figures match the spreadsheet, the quote is really in the contract, the policy says what the summary claims it says, the date is current. Completeness asks whether the output does everything it was asked to do: all twelve regions, all three questions, the risks section you requested, the caveat that the data stops in June. They fail independently, and an output can pass one while failing the other.
| The output is… | What it looks like | Why it slips through |
|---|---|---|
| Accurate but incomplete | Every figure is right, but two of the nine product lines are missing | What is there checks out, so nobody counts what isn’t |
| Complete but inaccurate | Every section is present; one margin is transposed (38% for 83%) | The structure matches the brief, so it “looks done” |
| Accurate and complete | All sections present, all figures traced, scope and caveats stated | This is the only one that is ready to send |
Anthropic’s own help centre is blunt about why this matters: Claude can write things that look correct but are mistaken, can present inaccurate quotes with confidence, can be confused about recent events because its training data has a cutoff, and can misinterpret sources even when it cites them. None of these failures announce themselves. A wrong number is formatted exactly like a right one.
Looks finished versus is finished
Signals that mislead
- Confident, fluent, professional tone
- Tidy headings, tables and bullet points
- Length — “it’s long, so it must be thorough”
- Claude says it is sure when you ask
Evidence that counts
- Each requirement in the brief is met
- Each input (file, row, region) is accounted for
- Each claim traces to a source you opened
- Calculations recomputed; gaps and assumptions stated
Checking completeness: count against the brief
Completeness is the easier check to do and the easier one to forget. The reliable method is to write your checklist from the request before you read the answer, so the answer cannot shape what you look for. Then tick it off.
- Requirements. List every instruction in your prompt: sections, questions, length, format, audience. Did each one happen?
- Inputs. If you gave Claude 14 files, 12 cost centres or 40 call notes, does the output account for all of them — or say which it could not use?
- Scope. Right time period, right regions, right version of the policy? A summary of the 2024 handbook is complete and wrong if you needed 2025.
- Caveats. Are limits stated — missing data, assumptions made, questions it could not answer? Silence is not the same as “nothing to flag”.
Some gaps come from the request, not the model. Anthropic’s prompting guidance notes that current Claude models follow instructions precisely and tend to do what was asked rather than volunteering extras; if you want risks, counter-arguments or next steps, ask for them. It compares Claude to a brilliant new employee who lacks context on your norms. An output that omits what you never asked for is a prompting problem; an output that omits what you did ask for is an evaluation finding.
Other gaps come from what Claude could actually read. The Claude help centre says that for PDFs of 100 pages or fewer Claude analyses both text and visual elements such as charts, but longer PDFs are processed as text only. For other document types such as Word files, Claude extracts text only and cannot read embedded images. Excel files need the code execution feature turned on. And in a very long conversation Claude may summarise earlier messages to keep going, so an instruction from an hour ago is worth restating if it matters.
Priya’s checklist, first draft
- Missing: All 12 cost centres consideredonly 10 appear; two large overspends absent
- Fails: Every variance over 5% listedfollows from the missing cost centres
- Fails: Figures match the exportWarehousing 7.2% shown; export gives 12.7%
- Check: A cause for each varianceFleet “fuel costs” is not in the data — a guess
- Passes: One page, CFO-ready format
Checking accuracy: where did each claim come from?
You cannot verify every sentence with the same effort, and you don’t need to. The practical move is to sort the claims that matter by where they came from, because each origin has its own check.
Tracing a claim to its check
- Your uploaded documentsFind the exact passageask Claude for the quote
- A calculationRecompute it yourselftotals, percentages, dates
- Web search or ResearchOpen the cited sourcedoes it say this, for this case?
- Claude’s general knowledgeVerify independentlymay be outdated or wrong
Citations deserve care. With web search on, Claude’s answers include citations and source links you can click, and Research runs many searches and returns a report with citations designed to be easy to check. But a citation tells you where Claude looked, not that it read the page correctly. The help centre explicitly advises reviewing the cited sources, because synthesis can misinterpret source material. The typical failure is not a fake link — it is a real page that says something slightly different: another year, another region, a forecast reported as a fact.
When you ask Claude in chat to cite page numbers or quote your own documents, treat those references as claims too. Anthropic’s developer documentation contrasts prompt-based citations, which it warns may be inaccurate, with a dedicated citations feature for applications. In the Claude apps, the check is simple: open the document and find the passage.
Making an answer checkable
Hard to checktext
Summarise the key terms of the
attached supplier contract.Built to be checkedtext
Summarise the attached supplier contract
for our procurement lead.
Cover, in this order: price and payment
terms, term and renewal, termination
rights, liability cap, SLAs.
For each term:
- quote the clause, with its number
- if the contract does not address it,
write "Not found in contract"
Use only the contract, not general
knowledge about typical contracts.
End with anything ambiguous or that
you were unsure about.The second prompt applies techniques from Anthropic’s guidance on reducing hallucinations: allow Claude to say it does not know, ground the answer in direct quotes, restrict it to the documents provided, and make every claim traceable. It also builds the completeness checklist into the request. You still check the quotes — but checking five clause numbers takes minutes, where checking an unanchored summary means re-reading the contract.
Match the depth of the check to the stakes
Evaluation takes time, and not every output deserves the same amount. Anthropic’s evaluation guidance makes the same point for builders: success criteria should be relevant to the use — citation accuracy matters enormously in some settings and far less in others. For everyday work, a simple scale helps.
| Stakes | Example | Reasonable check |
|---|---|---|
| Low — internal, easily undone | Brainstormed names for a team offsite | Read it; keep what’s useful |
| Medium — shared, informs work | Meeting summary sent to the project team | Checklist against the brief; spot-check key facts |
| High — external, financial, legal, or about people | Board figures, contract terms, HR policy, client proposals | Every figure recomputed, every claim traced, a qualified person reviews |
At the high end, review is not only good practice. Anthropic’s Usage Policy treats uses such as legal, healthcare, insurance, finance and employment as high-risk when outputs give advice or recommendations affecting individuals, and requires that a qualified professional review the content before it is finalised or disseminated. Lesson 2.4 covers when human review is required; for this objective, the point is that the stakes decide how far down the checklist you go. (This is general education about the policy, not legal advice.)
Traps the wrong answers are built from
| Tempting but wrong | Do this instead |
|---|---|
| Judging quality by tone, formatting or length | Check against the brief and the sources; polish is not evidence. |
| Asking “are you sure?” and accepting yes | Ask for the supporting quote or recompute the figure yourself. |
| Checking only what is there, not what is missing | Write the checklist from your request first, then count inputs and requirements. |
| Treating a working citation link as proof | Open the source and confirm it says this, for this year, place and case. |
| Applying the same light skim to every output | Scale the depth of checking to the cost of an error. |
You should now be able to
- Build a completeness checklist from the original request before reading the output.
- Distinguish accuracy failures from completeness failures and recognise each in a polished response.
- Trace important claims to their origin — your documents, a calculation, a cited web source or general knowledge — and apply the matching check.
- Rewrite a request so the answer is checkable: quotes, “not found” markers, labelled assumptions and stated scope.
- Explain why self-checking and citations reduce but do not replace verification.
- Match the depth of evaluation to the stakes of the output.