Rubric

Claude Certified Architect — Professional

CCAR-PChecked against the published guide on 11 September 2026

What the vendor publishes

Items
63
Unscored items
Not published by the vendor
Duration
120 minutes
Passing score
720 on a 100–1000 scale
Price
US$175
Validity
12 months

Multiple-choice and multiple-response items; each item states how many responses to select. Criterion-referenced: you pass by meeting a fixed standard, not by outperforming other candidates. Proctored by Pearson VUE, online or at a test centre. The score report gives pass/fail with the scaled score plus percent-correct by domain — section percentages are informational and do not determine the result.

Study the lessons Mock exams Soon

Domains and weightings

38 published task statements across 7 domains. Each links to its lesson once it is written.

The guide

What this exam actually tests

CCAR-P is about owning a system, not building a feature. Anthropic describes the credential as signalling that the holder can own or significantly contribute to the full lifecycle of a Claude-powered system — discovery, design, integration, evaluation, governance, handoff, iteration. The blueprint follows that arc literally, and it is why the exam feels different from Foundations even where the subject matter overlaps.

The clearest signal is in the weights. Integration is the largest domain at 19% — RAG pipeline design, retrieval strategy, authentication and authorization, observability at scale, choosing between MCP, API, CLI and agent-to-agent. And Stakeholder Communication & Lifecycle Management is 14%, the same weight as Governance. That is roughly one item in seven on structured discovery, communicating trade-offs, managing expectations and SLAs, and documenting an architecture for someone else to implement.

A technical exam that puts a seventh of its marks on communication is telling you something about the role it certifies. Candidates who prepare purely on the technical domains are preparing for about 85% of the paper.

Should you sit this or Foundations?

They are not sequential — there are no mandatory prerequisites for either, and the credential is awarded on exam performance alone. They test different jobs.

Anthropic’s recommended experience for CCAR-P is 3+ years in systems architecture or platform engineering, six or more months hands-on with Claude or comparable LLM systems, plus a foundation in software engineering practice and experience delivering systems from discovery through operationalisation. Foundations asks for roughly six months of hands-on work and no systems-architecture background.

The practical test: have you owned a system end to end — argued for the design, integrated it, evaluated it, and answered for it when it misbehaved in production? If yes, CCAR-P. If you have built well but not yet owned, Foundations is the honest entry point. The guide is explicit that the exam is aimed at a minimally qualified candidate with real-world deployment experience, and the blueprint has no way to be kind to someone without it.

The shape of the paper

Standalone items, not scenarios — this is the main structural difference from CCAR-F, which clusters its items under four scenarios. Sixty-three items in a hundred and twenty minutes is a shade under two minutes each, and without scenario setups to read, that average holds much more evenly across the paper.

Items are multiple-choice and multiple-response, each stating how many responses to select. The exam is criterion-referenced: you pass by meeting a fixed standard set through a formal standard-setting study, not by outperforming other candidates. There is no curve and no quota.

Your score report gives percent-correct by domain alongside the scaled score, and Anthropic states plainly that those section percentages are informational — the pass or fail rests on the total. Useful to know in the room: a domain going badly is not a failed exam.

From someone who sat it

This will carry a first-hand account of how the paper reads at pace — where the time goes, which domains reward deliberation and which punish it. It is being written from notes rather than published material, so it is held back until it can be said accurately.

The seven domains, and what each is really asking

Weights below are Anthropic’s own and are shown in full above this section. What follows is an interpretation for planning study time, not a claim about the item pool.

Integration — 19%

The largest domain, and the most concrete. Designing a RAG pipeline with sensible chunking and indexing; matching retrieval strategy to data shape and query pattern; spotting security gaps in authentication and authorization; choosing an integration mechanism and defending the choice; evaluating tool and agent configuration for capability bloat.

Two items in this domain reward production scars specifically: observability at scale, and progressive discovery versus loading everything into context. Both are the kind of thing you only get wrong once.

Solution Design & Architecture — 17%

Translating business problems into solutions, designing end to end including feedback loops, and selecting among workflow, agentic and augmented-LLM patterns. The distinctive objective is aligning a solution to business value pillars — efficiency, cost, productivity, performance SLAs. That framing is doing real work: it is asking whether you can justify an architecture in the terms the person funding it uses.

Evaluation, Testing & Optimization — 16%

Defining metrics across accuracy, latency, cost, safety and security; building evaluation datasets and mixed-methodology test frameworks; A/B testing; diagnosing prompt failure, hallucination and model mismatch as distinct problems with distinct fixes.

If you have never built an eval set, build one before sitting. It is the fastest domain to become competent in and the easiest to fake badly.

Governance, Safety & Risk Management — 14%

Guardrails and safety controls, failure modes of LLM systems, human-in-the-loop validation, named regulatory regimes (GDPR, HIPAA, FedRAMP), and ethical considerations — bias, fairness, transparency. Regulation appears by name in the blueprint, so treat it as examinable rather than as background colour.

Stakeholder Communication & Lifecycle Management — 14%

Structured discovery, communicating architectural decisions and trade-offs, managing feedback loops and expectations including SLAs, documenting for implementers, supporting each lifecycle phase from discovery through iteration.

The domain most often dismissed and, at 14%, the most expensive to dismiss. Prepare by rehearsing the explanation rather than the decision: given an architecture you chose, how would you justify it to a sceptical engineering lead, and separately to the person paying?

Claude Models, Prompting & Context Engineering — 13%

Model selection on trade-offs, system prompts and guardrails, zero-shot through chain-of-thought, context window and token management, and prompt reuse through caching, modular prompts and Skills. Smaller than candidates expect, because at this level prompting is treated as one component of a system rather than the subject itself.

Developer Productivity & Operational Enablement — 7%

The smallest domain: configuring Claude tooling for teams, improving developer workflows, supporting debugging and operational resolution. Four or five items. Worth an evening, not a week.

How to prepare

Anthropic’s own guidance is to build and operate at least one end-to-end Claude solution including RAG, evaluation and observability — and that single instruction covers the three largest domains at once. If you do nothing else, do that.

  • Build the end-to-end system. Retrieval, integration, evals, observability. Not a demo — something with a failure mode you had to fix.
  • Self-assess against each objective. The blueprint has 38 of them and they are published above. Score yourself honestly on each; the low ones are your plan.
  • Rehearse architectural decisions aloud. Model selection, integration protocol, security trade-offs. Say the justification out loud, because a seventh of the paper is about exactly that.
  • Work the sample questions in section 8 of the exam guide. Three published items with full rationale — the closest thing to a calibration you will get from the vendor.

Three things candidates get wrong

Preparing for CCAR-F, harder. Foundations is Claude Code, the Agent SDK, MCP and prompting. Professional is architecture, integration, evaluation, governance and communication. The overlap is smaller than the shared name suggests, and the items are not scenario-clustered.

Skipping the communication domain. Fourteen percent, and the easiest marks on the paper for anyone who has actually presented a design to stakeholders — which, given the three-years-of-architecture profile, is most candidates. Read the five objectives and recognise the job you already do.

Treating governance as the safety lecture. Named regulations, concrete guardrail decisions, and where a human belongs in the loop are all examinable. This is the domain where a candidate who has only worked on internal tools is most exposed.

What Rubric will have for this exam

Mock exams built to the seven published weightings, standalone-item format, every answer citing the documentation behind it and dated against the current blueprint. Results report per domain, matching the shape of the vendor’s own score report.

None of that exists yet. This page is the blueprint; the exams follow.

Where this came from

Every figure on this page was read from Anthropic’s published material, not from anyone’s exam. Blank fields are blank because the vendor does not publish them.

Source: anthropic-partners.skilljar.com

Vendors revise exams without much warning. If something here is out of date, tell us and it gets fixed and re-dated.

Claude Certified Architect — Professional (CCAR-P) — exam blueprint · Rubric