Two ways to present the same pilot
How credibility is lost
- “Claude makes us 80% faster” — a published average, quoted as our result
- “It’s basically accurate” — no definition, no measurement
- Limitations raised only when someone asks
- Benefits in hours saved, with no idea what the hours are for
How credibility is built
- “Our four-week pilot cut triage from 9 to 4 minutes per ticket”
- “One in six drafts needed a factual correction; here is the log”
- Limits stated up front, with the control that handles each
- Benefit named as the thing it buys: same-day response times
Say what kind of value you mean
“Productivity” is too vague to defend. Stakeholders fund one of four specific things, and naming which one you mean changes the conversation from an argument about AI into a discussion about a business outcome.
| Kind of value | What it means concretely | How you evidence it |
|---|---|---|
| Time | The same work takes fewer hours | Timed baseline before, same measure after |
| Capacity | More of the work gets done, or done sooner | Volume handled, backlog age, response times |
| Consistency | Four people produce the same shape of output | Rework rate, items returned by legal or QA |
| Risk reduction | Fewer things slip through | Errors caught at review, exceptions flagged earlier |
Time is the easiest to measure and the weakest to present on its own, because an executive’s next question is what happened to the freed time. The Advantage Solutions story Anthropic publishes handles this well: a compliance check that took 20 to 30 minutes on paper became a photograph and an automated check completed in minutes, and the company describes the result as redirecting more than 70,000 labour hours a year toward teammate training and floor presence. The hours went somewhere named. That is the sentence a board hears.
Your own numbers beat everyone else’s
There is a hierarchy of evidence, and it is worth being explicit about which rung you are standing on. Your own measured pilot on your own work is the strongest. A published case study from a comparable organisation is next, useful for showing something is possible but never as a prediction for you. Research averages are context, not forecast. A vendor claim, including one from Anthropic, is the weakest thing you can put in front of a finance director.
Anthropic’s own research is a good demonstration of why. In November 2025 the company published an analysis estimating productivity gains from around 100,000 Claude.ai conversations, in which Claude estimated how long each task would have taken without AI and how long it took with it. The headline is striking — time savings concentrated in the 50 to 95 per cent range and clustering around 80 per cent, on tasks with a median estimated duration of 1.4 hours. The paper’s own caveats are the more useful part for anyone presenting to stakeholders.
- The estimates are the model’s, and the authors state they lack real-world data to validate them for most tasks.
- The analysis does not count time users spend on the task outside the conversation — checking, correcting and refining the output.
- The data comes only from people already using Claude, and from tasks they expected it to handle well, which is a selection effect the authors name.
- A referenced trial on end-to-end software features saw no time savings at all, which is a reminder that task-level gains need not survive at the level of whole pieces of work.
Which limitations to state, and how
Limitations land better when each one arrives attached to the control that handles it. Anthropic’s own help centre is candid about the failure modes: Claude can write things that look correct but are mistaken; it may be confused about current events because its training data has a cutoff; it can misinterpret sources it has searched; and it can produce quotations that sound authoritative but are not grounded in fact. The guidance to users is equally plain — do not treat Claude as a singular source of truth, and scrutinise high-stakes advice carefully.
| Limitation | How to say it | The control you pair it with |
|---|---|---|
| Confident errors | “It writes wrong answers in exactly the same tone as right ones.” | Review before consequence; trace claims to sources |
| Training cutoff | “It does not know what happened after its training data ends.” | Web search or Research for anything current |
| Misread sources | “A citation shows where it looked, not that it read the page correctly.” | Open the cited source for claims that matter |
| No access by default | “It cannot see our systems unless we connect them or paste the data.” | Connectors, uploads, or figures from the system of record |
| Input limits | “Very long documents are read as text; some content is not read at all.” | Know the file limits; check what was actually covered |
| Availability differs | “Some features need a paid plan or an admin to enable them.” | Confirm plan and organisation settings before promising |
Two of those are worth getting exactly right, because stakeholders will test them. On file handling, Anthropic’s documentation states that PDFs of up to 100 pages are analysed for both text and visual elements, that longer PDFs — up to a 1,000-page limit — are processed as text only with visual elements not analysed, and that for formats such as Word files Claude extracts text only and cannot interpret embedded images. On availability, Research requires a paid plan and web search to be enabled, and it can consume usage limits faster than ordinary chats because it retrieves many sources. Saying this before a stakeholder discovers it is the difference between a caveat and a complaint.
The same update, rewritten for a sceptic
Reads as marketingtext
AI update: the Claude rollout has been a
huge success. The team is far more
productive and the quality of our work
has improved dramatically. Industry
research shows AI can cut task time by
up to 80%. We recommend expanding to
all departments next quarter.Reads as evidencetext
Claude pilot, weeks 1-4 (8 staff,
380 items of client correspondence).
Drafting time: 22 min to 9 min per item.
Same-day turnaround: 61% to 88%.
Factual errors caught at review: 1 in 9
- all corrected before sending.
Limits: no access to the policy system;
nothing after the training cutoff;
confident when wrong, so review stays.
Cost: licences, ~90 min training each,
6 weeks building templates.
Ask: 8 more weeks, two more letter
types, same measures, decide after.The questions you will actually be asked
Most stakeholder conversations turn on four questions, and each has an honest answer that is better than the evasive one.
- “Is it accurate?” Not reliably enough to go unreviewed. Say what your review catches and how often — a measured error rate is a stronger answer than a reassurance.
- “Will this replace jobs?” Answer the question that was asked, and say what is actually changing: which tasks move, which stay, and what the freed time is for. Vagueness here is heard as evasion.
- “Where does our data go?” Answer from your organisation’s plan, contract and policy, not from memory. If you do not know, say so and route it to whoever owns data protection. This is Domain 6 territory, and 6.2 covers it.
- “What does it cost?” Include the setup work, the training, the review time that does not disappear, and the ongoing ownership of templates and instructions — not just licences.
Is this claim ready to present?
- Passes: The number comes from our work, not a published average380 items over four weeks
- Passes: The measure is defined the same way before and afterminutes per item of correspondence
- Passes: The scope of the claim is statedone correspondence type, eight staff
- Passes: The error rate is disclosed1 in 9, all caught at review
- Check: The limitations are stated with their controlsdata-handling answer still with the DPO
- Fails: The full cost is includedtraining and template-building time not counted
- Passes: The ask is specific and time-boxedeight weeks, two letter types
Traps the wrong answers are built from
| Tempting but wrong | Do this instead |
|---|---|
| Quoting a published productivity figure as your own expected result | Use published work to show a direction is plausible; commit only to your measured pilot. |
| Presenting a pilot with no error data | Disclose the error rate and the control that catches it; it is what makes the rest credible. |
| Raising limitations only when challenged | State each limit up front, paired with the control that handles it. |
| Answering a data-protection question from memory | Answer from your plan, contract and policy, or route it to the owner and come back. |
| Counting licences as the cost | Include setup, training, ongoing review and who maintains the templates and instructions. |
You should now be able to
- Name which kind of value — time, capacity, consistency or risk — a claim is about, and evidence it accordingly.
- Rank evidence from your own pilot down to vendor claims, and say which rung you are on.
- State Claude’s limitations accurately, including training cutoff, confident errors, source misreading and file handling.
- Pair every limitation with the control that handles it.
- Disclose error rates and full costs in a stakeholder update.
- Answer the accuracy, jobs, data and cost questions honestly rather than evasively.