Rubric
Contents — domains, guide and mocks

When a human must review

CCAO-F 2.410 min read · checked 21 September 2026

Task statementDetermine when human review or additional verification is required

Who needs to see this first?

What will this output be used for?
  • Just me; easy to undo
    Your own checkaccuracy and completeness pass
  • Others will rely on it
    Verify key factssources checked; second reader
  • Affects a person’s rights, money, health
    Expert sign-offqualified reviewer, before release
  • Claude will take an action
    Approve before it runsespecially if irreversible
Where the output goes and who it affects set the level of review — not how good the draft looks.

What raises the bar

Every output gets the basic evaluation from 2.1. The question here is when that is not enough — when you need additional verification against sources (the techniques in 2.3), or a second person, or a person with specific expertise. A handful of factors decide it, and any one of them can raise the bar on its own.

FactorLower barHigher bar
AudienceOnly you, or your immediate teamCustomers, regulators, the public, the board
ReversibilityA draft you can revise tomorrowA sent email, a signed contract, a payment, a deletion
Who is affectedNo one in particularSpecific individuals — their job, credit, care, claim
DomainBrainstorming, formatting, internal notesLegal, medical, financial, employment, insurance, housing
Where the content came fromYour own documents, quotedClaude’s general knowledge, recent events, uncited claims
What your checks foundNothing unusualAny hallucination, inconsistency or bias signal (2.2)
VolumeOne output, read in fullHundreds of outputs no one reads individually

The last row is easy to miss. Reviewing one email is simple; a process that generates hundreds of customer replies or candidate assessments needs its own review design, because no single person sees every output. The more automated the flow, the more deliberately the human checkpoints have to be placed.

Levels of oversight

More oversight further down

  1. Self-checkbrief met, claims plausible, nothing missing
  2. Additional verificationkey claims traced to primary sources
  3. Second reviewera colleague reads before it goes out
  4. Qualified professionallegal, clinical, financial or HR sign-off
  5. Human decidesClaude informs; a person makes the call
Move down the stack as the stakes rise. The bottom level is not more checking — it is keeping the decision itself with a person.

Where the rules set the bar for you

In some areas the decision is not yours to make. Anthropic’s Usage Policy lists high-risk use cases — including legal, healthcare, insurance, financial, employment and housing, academic testing, and media uses — where outputs that give advice or recommendations, or feed subjective decisions directly affecting individuals, require two safeguards. First, human-in-the-loop: a qualified professional in that field reviews the content or decision before it is disseminated or finalised. Second, disclosure: people are told that AI was used to help produce it.

Data protection law can reach the same place. Article 22 of the EU’s GDPR gives people the right not to be subject to a decision based solely on automated processing that produces legal effects or similarly significant effects on them, with limited exceptions; where such decisions are allowed, safeguards include the right to obtain human intervention and to contest the decision. Anthropic’s own research on discrimination in model decisions makes the same point from the other side: it does not endorse using language models to make automated decisions in high-risk areas such as lending and hiring. Your organisation’s AI policy may add more — following it is covered in 6.3.

Review that actually counts

A review is only as good as the conditions it happens under. A busy manager clicking “approve” on forty AI-drafted letters in ten minutes has not reviewed them in any meaningful sense. Anthropic’s documentation is clear that techniques for reducing hallucinations lower the error rate but do not eliminate it, so a reviewer has to be positioned to catch what remains.

Is this review meaningful?

  • Passes: Reviewer is qualified for the subjectclinician for medical content; lawyer for legal
  • Passes: Happens before the output is sent or finalnot after complaints arrive
  • Check: Reviewer has the sources, not just the draftwithout them they can only judge the prose
  • Fails: Enough time for the volumeforty letters in ten minutes is not review
  • Passes: Reviewer can reject or change it
  • Missing: What was checked is recordedneeded if the decision is questioned later
A review that fails any of these is closer to a formality than a safeguard.

When Claude acts, not just writes

Some Claude products can take actions: Cowork, for example, can work with your files, browser and apps, and depending on the permissions you grant it can delete files, send messages or make purchases. Anthropic’s safety guidance for Cowork names two risks: destructive actions, and prompt injection, where instructions hidden in content Claude reads — a web page, a document, an email — try to redirect it. It recommends limiting access to what the task needs, monitoring tasks, using manual approval for high-stakes work, and remembering that you remain responsible for what Claude does on your behalf.

Traps the wrong answers are built from

Tempting but wrongDo this instead
Adding an AI disclaimer instead of reviewingDisclose and review; disclosure does not make content correct.
Reviewing after the output has gone outReview before it is sent, published or acted on.
Any available colleague as the reviewerWhere it affects rights, health, money or legal position, use a qualified professional.
The same review for every outputScale oversight to audience, reversibility, who is affected and domain.
Letting Claude make irreversible changes unattendedLimit access, use manual approval, and verify through an independent channel.

You should now be able to

  • Identify the factors — audience, reversibility, who is affected, domain, provenance, volume — that raise the level of review.
  • Choose between self-check, additional verification, second reviewer, qualified sign-off and human decision.
  • Recognise the Usage Policy’s high-risk categories and its human-in-the-loop and disclosure safeguards.
  • Explain why solely automated decisions about individuals need meaningful human involvement.
  • Tell a meaningful review from a formality.
  • Place approval before Claude takes consequential actions.

Practice questions

Original questions written for this lesson, in the exam’s style. Answer first, then open the reasoning — every option is explained, including why the wrong ones are tempting.

  1. Question 1

    Which of these outputs most clearly requires review by a qualified professional before it is sent?

    1. AA list of team-building ideas for an internal offsite.
    2. BA summary of this morning’s project meeting for the team.
    3. CA letter telling a policyholder why their insurance claim is denied.
    4. DA first outline of headings for a company blog post.
    Show answer and reasoning
    1. AIncorrect. Internal, low-stakes and easily changed; a self-check is enough.
    2. BIncorrect. Colleagues will rely on it, so a quick check of decisions and owners is sensible, but no specialist is needed.
    3. CCorrect. It affects an individual’s money and rights in insurance, a high-risk category where a qualified person must review before release.
    4. DIncorrect. A draft structure that will be rewritten; nothing is disseminated yet.
  2. Question 2

    A consultancy wants to speed up client reporting. It proposes adding “This report was generated with AI and may contain errors” to each report and dropping the partner review step.

    What is the best response?

    1. AAccept it, because the disclaimer shifts responsibility to the client.
    2. BKeep the review; the disclaimer is disclosure, not a check on accuracy.
    3. CAccept it, and review a sample of reports after they are sent.
    4. DReplace the partner review with a request for Claude to self-check.
    Show answer and reasoning
    1. AIncorrect. A disclaimer does not make the content accurate, and the firm remains responsible for what it sends.
    2. BCorrect. Disclosure and review do different jobs. Client-facing advice still needs a qualified person to check it before release.
    3. CIncorrect. After-the-fact sampling is monitoring; errors reach clients before anyone checks.
    4. DIncorrect. Self-checking is not independent and is no substitute for a qualified reviewer.
  3. Question 3

    An operations lead wants Cowork to read supplier emails and update bank details in the payments spreadsheet overnight, unattended, so it is ready for the morning payment run.

    What is the most appropriate change to this plan?

    1. ALet it run, since it only reads the company’s own inbox.
    2. BRun it as planned and review the spreadsheet after the payment run.
    3. CHave it propose changes for approval, confirmed by phone on file.
    4. DAdd an instruction telling Claude to ignore suspicious emails.
    Show answer and reasoning
    1. AIncorrect. The emails come from outside the company; untrusted content plus a consequential change is the risky combination.
    2. BIncorrect. Reviewing after payments are made is too late to prevent a fraudulent change.
    3. CCorrect. Manual approval before the change and verification through an independent channel address both prompt injection and payment fraud.
    4. DIncorrect. A prompt instruction may help, but it is not a control; hidden instructions are designed to look legitimate.
  4. Question 4

    An HR team in the EU sets up a workflow where Claude scores job applications and automatically rejects everyone below a threshold. No one looks at the rejected applications.

    What is the main problem with this design?

    1. ACandidates may be rejected by a solely automated decision with no human involvement.
    2. BScoring many applications in one chat may exceed the context window.
    3. CAutomatically generated rejection emails may sound impersonal to candidates.
    4. DClaude may be slower than a recruiter at reading each application.
    Show answer and reasoning
    1. ACorrect. Hiring decisions significantly affect individuals; GDPR Article 22 and Anthropic’s policy both point to meaningful human involvement.
    2. BIncorrect. A practical concern that can be designed around; it is not the core issue with the workflow.
    3. CIncorrect. Tone matters, but it is minor next to excluding people with no human review.
    4. DIncorrect. Speed is not the issue; the design removes human judgment from a decision about people.

Sources

Drafted with AI assistance and checked against the sources above; expert review is in progress. Spotted an error? Tell us and it gets fixed, dated and listed on how this is written.