Rubric
Contents — domains, guide and mocks

Checking outputs for accuracy and completeness

CCAO-F 2.115 min read · checked 21 September 2026

Task statementEvaluate Claude-generated outputs for accuracy and completeness

The evaluation pass

  1. Restate the briefwhat you asked, for whom, from which inputs
  2. Is it complete?every requirement and input covered
  3. Is it accurate?trace each claim to its source
  4. Weigh the stakeshow costly is an error here?
  5. Fix or ask againcorrect it yourself, or re-prompt precisely

Re-check the revised output the same way — a fix can introduce a new error

Completeness is checked against your request; accuracy is checked against the sources. Neither is checked against how confident the answer sounds.

Two different questions

Accuracy asks whether what the output says is true: the figures match the spreadsheet, the quote is really in the contract, the policy says what the summary claims it says, the date is current. Completeness asks whether the output does everything it was asked to do: all twelve regions, all three questions, the risks section you requested, the caveat that the data stops in June. They fail independently, and an output can pass one while failing the other.

The output is…What it looks likeWhy it slips through
Accurate but incompleteEvery figure is right, but two of the nine product lines are missingWhat is there checks out, so nobody counts what isn’t
Complete but inaccurateEvery section is present; one margin is transposed (38% for 83%)The structure matches the brief, so it “looks done”
Accurate and completeAll sections present, all figures traced, scope and caveats statedThis is the only one that is ready to send

Anthropic’s own help centre is blunt about why this matters: Claude can write things that look correct but are mistaken, can present inaccurate quotes with confidence, can be confused about recent events because its training data has a cutoff, and can misinterpret sources even when it cites them. None of these failures announce themselves. A wrong number is formatted exactly like a right one.

Looks finished versus is finished

Signals that mislead

  • Confident, fluent, professional tone
  • Tidy headings, tables and bullet points
  • Length — “it’s long, so it must be thorough”
  • Claude says it is sure when you ask

Evidence that counts

  • Each requirement in the brief is met
  • Each input (file, row, region) is accounted for
  • Each claim traces to a source you opened
  • Calculations recomputed; gaps and assumptions stated
Everything on the left is a property of the writing. Everything on the right is a property of the evidence. The exam rewards the right-hand column.

Checking completeness: count against the brief

Completeness is the easier check to do and the easier one to forget. The reliable method is to write your checklist from the request before you read the answer, so the answer cannot shape what you look for. Then tick it off.

  1. Requirements. List every instruction in your prompt: sections, questions, length, format, audience. Did each one happen?
  2. Inputs. If you gave Claude 14 files, 12 cost centres or 40 call notes, does the output account for all of them — or say which it could not use?
  3. Scope. Right time period, right regions, right version of the policy? A summary of the 2024 handbook is complete and wrong if you needed 2025.
  4. Caveats. Are limits stated — missing data, assumptions made, questions it could not answer? Silence is not the same as “nothing to flag”.

Some gaps come from the request, not the model. Anthropic’s prompting guidance notes that current Claude models follow instructions precisely and tend to do what was asked rather than volunteering extras; if you want risks, counter-arguments or next steps, ask for them. It compares Claude to a brilliant new employee who lacks context on your norms. An output that omits what you never asked for is a prompting problem; an output that omits what you did ask for is an evaluation finding.

Other gaps come from what Claude could actually read. The Claude help centre says that for PDFs of 100 pages or fewer Claude analyses both text and visual elements such as charts, but longer PDFs are processed as text only. For other document types such as Word files, Claude extracts text only and cannot read embedded images. Excel files need the code execution feature turned on. And in a very long conversation Claude may summarise earlier messages to keep going, so an instruction from an hour ago is worth restating if it matters.

Priya’s checklist, first draft

  • Missing: All 12 cost centres consideredonly 10 appear; two large overspends absent
  • Fails: Every variance over 5% listedfollows from the missing cost centres
  • Fails: Figures match the exportWarehousing 7.2% shown; export gives 12.7%
  • Check: A cause for each varianceFleet “fuel costs” is not in the data — a guess
  • Passes: One page, CFO-ready format
The checklist was written from the brief before she read the output. Two items fail and one is missing — none of which the formatting gave away.

Checking accuracy: where did each claim come from?

You cannot verify every sentence with the same effort, and you don’t need to. The practical move is to sort the claims that matter by where they came from, because each origin has its own check.

Tracing a claim to its check

Where did this claim come from?
  • Your uploaded documents
    Find the exact passageask Claude for the quote
  • A calculation
    Recompute it yourselftotals, percentages, dates
  • Web search or Research
    Open the cited sourcedoes it say this, for this case?
  • Claude’s general knowledge
    Verify independentlymay be outdated or wrong
Every important claim should land on one of these four branches. A claim you cannot trace to any of them is not a finding — it is a guess to verify or delete.

Citations deserve care. With web search on, Claude’s answers include citations and source links you can click, and Research runs many searches and returns a report with citations designed to be easy to check. But a citation tells you where Claude looked, not that it read the page correctly. The help centre explicitly advises reviewing the cited sources, because synthesis can misinterpret source material. The typical failure is not a fake link — it is a real page that says something slightly different: another year, another region, a forecast reported as a fact.

When you ask Claude in chat to cite page numbers or quote your own documents, treat those references as claims too. Anthropic’s developer documentation contrasts prompt-based citations, which it warns may be inaccurate, with a dedicated citations feature for applications. In the Claude apps, the check is simple: open the document and find the passage.

Making an answer checkable

Hard to checktext

Summarise the key terms of the
attached supplier contract.

Built to be checkedtext

Summarise the attached supplier contract
for our procurement lead.

Cover, in this order: price and payment
terms, term and renewal, termination
rights, liability cap, SLAs.

For each term:
- quote the clause, with its number
- if the contract does not address it,
  write "Not found in contract"
Use only the contract, not general
knowledge about typical contracts.

End with anything ambiguous or that
you were unsure about.
The right-hand prompt names every term it expects, asks for the clause behind each one, and makes a missing term show up as “Not found” instead of vanishing.

The second prompt applies techniques from Anthropic’s guidance on reducing hallucinations: allow Claude to say it does not know, ground the answer in direct quotes, restrict it to the documents provided, and make every claim traceable. It also builds the completeness checklist into the request. You still check the quotes — but checking five clause numbers takes minutes, where checking an unanchored summary means re-reading the contract.

Match the depth of the check to the stakes

Evaluation takes time, and not every output deserves the same amount. Anthropic’s evaluation guidance makes the same point for builders: success criteria should be relevant to the use — citation accuracy matters enormously in some settings and far less in others. For everyday work, a simple scale helps.

StakesExampleReasonable check
Low — internal, easily undoneBrainstormed names for a team offsiteRead it; keep what’s useful
Medium — shared, informs workMeeting summary sent to the project teamChecklist against the brief; spot-check key facts
High — external, financial, legal, or about peopleBoard figures, contract terms, HR policy, client proposalsEvery figure recomputed, every claim traced, a qualified person reviews

At the high end, review is not only good practice. Anthropic’s Usage Policy treats uses such as legal, healthcare, insurance, finance and employment as high-risk when outputs give advice or recommendations affecting individuals, and requires that a qualified professional review the content before it is finalised or disseminated. Lesson 2.4 covers when human review is required; for this objective, the point is that the stakes decide how far down the checklist you go. (This is general education about the policy, not legal advice.)

Traps the wrong answers are built from

Tempting but wrongDo this instead
Judging quality by tone, formatting or lengthCheck against the brief and the sources; polish is not evidence.
Asking “are you sure?” and accepting yesAsk for the supporting quote or recompute the figure yourself.
Checking only what is there, not what is missingWrite the checklist from your request first, then count inputs and requirements.
Treating a working citation link as proofOpen the source and confirm it says this, for this year, place and case.
Applying the same light skim to every outputScale the depth of checking to the cost of an error.

You should now be able to

  • Build a completeness checklist from the original request before reading the output.
  • Distinguish accuracy failures from completeness failures and recognise each in a polished response.
  • Trace important claims to their origin — your documents, a calculation, a cited web source or general knowledge — and apply the matching check.
  • Rewrite a request so the answer is checkable: quotes, “not found” markers, labelled assumptions and stated scope.
  • Explain why self-checking and citations reduce but do not replace verification.
  • Match the depth of evaluation to the stakes of the output.

Practice questions

Original questions written for this lesson, in the exam’s style. Answer first, then open the reasoning — every option is explained, including why the wrong ones are tempting.

  1. Question 1

    A sales operations manager asks Claude to summarise 40 customer call notes into the top objections by region, for a VP review tomorrow. The response is well organised, with five themes and a confident tone.

    What is the most appropriate next step before sending it?

    1. AAsk Claude whether it is confident the summary is accurate.
    2. BCheck that every region and note is accounted for, then trace key themes back to the notes.
    3. CRegenerate the summary twice and send the version that reads best.
    4. DProofread it for tone and formatting so it suits a VP audience.
    Show answer and reasoning
    1. AIncorrect. A confident reply adds no evidence; the model can be confidently wrong, and this checks nothing against the notes.
    2. BCorrect. This checks completeness against the brief and accuracy against the source — the two halves of evaluation.
    3. CIncorrect. Comparing runs can reveal inconsistencies, but choosing by readability rewards polish, not correctness.
    4. DIncorrect. Tone matters for the audience, but it is not an accuracy or completeness check.
  2. Question 2

    An analyst uploads a 180-page annual report, full of charts, to a Claude chat and asks for commentary on year-on-year changes in operating margin. She plans to use the figures in a client presentation.

    Which TWO actions most improve confidence in the accuracy and completeness of the figures? (Select 2.)

    1. AAsk Claude for a longer, more detailed response so nothing is left out.
    2. BAsk Claude to quote the passage behind each figure and flag any it cannot find.
    3. CAccept the figures, since the report was uploaded and Claude has read it.
    4. DRecompute the year-on-year changes herself from the report’s tables.
    5. ETurn on web search so Claude can confirm the figures online.
    6. FAsk Claude to rate its confidence in each figure from 1 to 10.
    Show answer and reasoning
    1. AIncorrect. Length is not completeness; a longer answer can omit the same items and add more to check.
    2. BCorrect. Quotes make each figure traceable, and a “cannot find” flag turns silent gaps into visible ones — important here because a PDF this long is read as text only, so chart-only figures may be missed.
    3. CIncorrect. Uploading does not guarantee every element was read or interpreted correctly; long PDFs are processed as text only.
    4. DCorrect. Calculations are claims too; recomputing is the direct check for derived figures.
    5. EIncorrect. It introduces other sources that may differ from the report; the check belongs against the document itself.
    6. FIncorrect. Self-rated confidence is not evidence and does not trace any figure to its source.
  3. Question 3

    Using web search, Claude tells a marketing lead that “62% of mid-market buyers now prefer self-serve purchasing”, with a citation. She clicks the link: it is a real industry report on the right topic, but the 62% figure refers to enterprise buyers in a different year.

    What does this situation best illustrate?

    1. AThe claim is verified, because the citation links to a real report.
    2. BWeb search should be avoided for statistics in business content.
    3. CA citation shows where Claude looked; the claim still needs checking against the page.
    4. DThis is not an error, because hallucinations only involve made-up sources.
    Show answer and reasoning
    1. AIncorrect. A working link shows where Claude looked, not that the claim matches what the page says.
    2. BIncorrect. Web search is useful for current information; the lesson is to check citations, not to avoid them.
    3. CCorrect. Synthesis can misread a real source — the wrong segment and year here — so opening the source is the check.
    4. DIncorrect. Misrepresenting a real source is still inaccurate output, and it is the more common form with citations.
  4. Question 4

    An HR partner asks Claude, in a Project containing five office handbooks, to compare parental leave policies. The table covers four offices; the fifth is silently missing. She wants to prevent this next time.

    Which change to her request best addresses the problem?

    1. AName all five offices and require a row for each, marked “not found” if absent.
    2. BAdd “Please be thorough and don’t miss anything” to the request.
    3. CAdd several paragraphs of background on why the comparison matters.
    4. DSwitch to the most capable model for all HR comparison tasks.
    Show answer and reasoning
    1. ACorrect. An explicit list defines completeness, and the “not found” rule makes any gap visible instead of silent.
    2. BIncorrect. A vague instruction gives no checklist; Claude cannot tell which items “anything” refers to.
    3. CIncorrect. Context can help quality, but it does not define which items must appear.
    4. DIncorrect. A stronger model may help, but without a defined list you still cannot see what was omitted.

Sources

Drafted with AI assistance and checked against the sources above; expert review is in progress. Spotted an error? Tell us and it gets fixed, dated and listed on how this is written.