🛠️

🤖 Anthropic

Free Claude Certified Developer - Foundations (CCDV-F) Study Resources

Build and ship Claude applications the way the exam tests it — agents and workflows, the Claude API and SDKs, Claude Code, debugging, model selection and cost, prompt and context engineering, security and guardrails, and tools and MCP servers, mapped 1:1 to the official CCDV-F exam guide. The exam itself is currently available only to organizations in the Claude Partner Network.

503
Practice Questions
17
Study Notes
222
Flashcards
Start Studying — Free1 learners studying this hive

Claude Certified Developer - Foundations (CCDV-F) Study Notes & Guides

17 AI-generated study notes covering the full Claude Certified Developer - Foundations (CCDV-F) curriculum. Showing 10 complete guides below.

Hands-on Lab1,954 words

Agents and Workflows — practice exercise

Read full article

Agents and Workflows — practice exercise

Difficulty: intermediate · Estimated duration: 60–90 minutes

You will design and write the core of a small agent system. The work is code and design written in your editor. No stage needs an API key or spends any credit. You compare what you write with the reference solution at the end. One optional step lets you run your loop against the live API; if you take it, it uses your own API key and your own credit.

The scenario is invented for practice. You are a developer on an accounts-receivable team. Customers dispute invoices by email, and the team wants Claude to help resolve the disputes.

StageWhat you practiseMinutes
1Workflow, agent or orchestrator, and subagents15
2Choosing a surface and writing the agent loop20
3Hosting the agent and guarding it with a hook15
4Tool contract, memory and workflow patterns20

Stage 1 — Choose the architecture

Skills: CCDVF-U1.T1.LO1.S1, CCDVF-U1.T1.LO1.S2, CCDVF-U1.T1.LO1.S3 Minutes: 15

The team lists three features:

  • A. Each dispute email is classified as pricing, delivery or duplicate charge, and a template reply is drafted for that category. The steps never change.
  • B. For a disputed invoice, find out what went wrong. Depending on the dispute this may mean checking the order system, the delivery tracker, the payment processor, or several of them. Nobody can list the checks in advance.
  • C. While working on B, the agent sometimes has to read a customer’s full payment history, which runs to thousands of lines, only to find two or three relevant transactions.
  1. For each of A and B, write one sentence: workflow or agent (or orchestrator-workers), and why.
  2. B could instead be built as orchestrator-workers. Write two sentences on what the orchestrator would decide, and what each worker would do.
  3. For C, describe a subagent: what it reads, what tools it needs, and what it returns to the main conversation.
  4. Before B touches real customer accounts, scope its tools and permissions: name the tools it may only read with, and the one action it must never be able to take on its own.

Stage 2 — Pick the surface and write the loop

Skills: CCDVF-U1.T2.LO2.S1, CCDVF-U1.T2.LO2.S2 Minutes: 20

The team runs Python services on its own servers and wants full control over each call.

  1. Choose a surface: the Agent SDK, the Client SDK with your own loop, or Managed Agents (beta). Justify it in two sentences.
  2. Write the agentic loop in Python for feature B, using the Client SDK. Assume a client, a TOOLS list, and a function run_tool(block) that executes one tool call and returns its output as a string, raising an exception if the tool fails. Your loop must:
    • keep one running message list;
    • stop when stop_reason is no longer "tool_use";
    • answer every tool_use block in a response, in one user message;
    • return a failing tool’s error with is_error: True instead of raising;
    • stop after 15 turns even if Claude has not finished, and say what happens then.
  3. (Optional; uses your own API key and credit.) Run the loop against one stub tool that returns a fixed invoice record, and check the loop exits cleanly.

Stage 3 — Host it and guard it

Skills: CCDVF-U1.T2.LO2.S3, CCDVF-U1.T2.LO2.S4 Minutes: 15

The team is now considering moving feature B to the Agent SDK in containers.

  1. Disputes now arrive as a continuous stream from a busy shared inbox, all day. Choose a session pattern from the hosting guide and justify it.
  2. List what agent state would live on the container’s disk, and what happens to it on a scale-down.
  3. Sketch a Python PreToolUse hook callback that denies any Bash command that updates the ledger table, with a reason the agent can read. Make the check robust to how people type SQL.

Stage 4 — Patterns, tools and memory

Skills: CCDVF-U1.T3.LO3.S1, CCDVF-U1.T3.LO3.S2, CCDVF-U1.T3.LO3.S3 Minutes: 20

  1. Claude needs to look up an invoice by its number before it can reason about a dispute. Write the tool definition: its name, a description Claude can act on, and its input schema.
  2. The agent should remember each customer’s past disputes between sessions, using the memory tool. Write handle(inp, store), the part of your memory handler that dispatches each memory-tool command (view, create, str_replace, insert, delete, rename) to a store object. Assume store already maps paths onto your storage and checks them. Say what your function returns when a command fails.
  3. Name the workflow pattern for each, in one line:
    • (a) refunds above a set amount get three independent assessments, and a refund goes ahead only if at least two agree;
    • (b) each proposed refund is checked at the same time by three separate calls, one for policy, one for fraud signals and one for the customer’s history, and the findings are combined;
    • (c) dispute emails arrive in five languages, and each is best handled by a prompt written for that language.

Acceptance checks

StageYou are done when
1A and B each have a design and a reason; C has a subagent with narrow tools and a summary output; B’s tools are scoped, with one action kept for a person
2Your loop meets all five rules, handles the cap, and your surface choice names who runs the loop
3Your pattern matches the traffic; your state list names what a scale-down loses; your hook denies with a reason, and its check survives case and spacing
4The tool has a clear description and schema; handle covers all six commands and answers a failure with a message, not an exception; the three patterns are named

Reference solution

Stage 1. A is a workflow: the steps are known and fixed, so predefined code paths give predictability and consistency. Classification into three categories followed by a template is also a natural fit for routing.

B needs an agent or orchestrator-workers: the checks depend on the dispute, so they cannot be hardcoded. As orchestrator-workers, the orchestrator reads the dispute and decides which systems to check. Each worker runs one check, such as querying the delivery tracker, and the orchestrator then synthesizes the findings.

For C, a subagent reads the payment history in its own context window, with read-only access to the payments tool. It returns only the relevant transactions as a short summary, so the history never floods the main conversation.

Scoping for B: read-only access to the order system, the delivery tracker and the payment processor, because investigating a dispute only needs to look. Issuing a refund or credit is not a tool it has: it recommends one, with its evidence, and a person carries it out. Scoping tools this way limits what a wrong step can do, whatever the model decides.

Stage 2. The Client SDK with your own loop fits a team that wants full control and runs its own Python services: with the Client SDK you write the tool loop yourself. The Agent SDK would also run in their process, but it brings Claude Code’s own tools and configuration, which this task does not need. A loop that meets the five rules:

python
USE = "tool_use" ask = client.messages.create def result(b): res = {"type": "tool_result", "tool_use_id": b.id} try: res["content"] = run_tool(b) except Exception as err: res["content"] = str(err) res["is_error"] = True return res msgs = [{"role": "user", "content": dispute}] for _ in range(15): # turn cap r = ask(model=MODEL, max_tokens=1024, tools=TOOLS, messages=msgs) msgs.append({ "role": "assistant", "content": r.content}) if r.stop_reason != USE: break out = [result(b) for b in r.content if b.type == USE] msgs.append({"role": "user", "content": out})

Each tool_result carries the tool_use_id of the call it answers. All results for one response go back in one user message, and a failing tool becomes an is_error result that Claude can react to. It can retry with corrected input, ask for clarification, or explain the limitation.

After the loop, check r.stop_reason:

  • end_turn is a final answer.
  • max_tokens or refusal need their own handling.
  • If the loop ended because it hit the 15-turn cap, stop_reason is still "tool_use": the last tool results were appended but never sent. Treat that as an unfinished run, and log it or hand it to a person rather than presenting it as an answer.

Stage 3. Long-running sessions fit a continuous stream: persistent container instances serve ongoing work, and the guide lists high-volume message streams among their uses.

By default the container’s disk holds the session transcripts, CLAUDE.md memory files and working-directory artifacts. None of them survive a restart, scale-down or move to another node, so anything needed later must go to storage outside the container.

A hook callback, following the shape in the hooks guide:

python
import re LEDGER = re.compile( r"\bupdate\s+ledger\b", re.I) async def block_ledger( data, tool_use_id, ctx): tool = data["tool_input"] cmd = tool.get("command", "") if not LEDGER.search(cmd): return {} out = dict( hookEventName= data["hook_event_name"], permissionDecision="deny", permissionDecisionReason=( "Ledger updates need " "finance review.")) return {"hookSpecificOutput": out}

A plain substring test such as "UPDATE ledger" in cmd misses update ledger or UPDATE ledger with two spaces. The regular expression ignores case and allows any whitespace. The reason string matters too: the agent reads it, so say what to do instead of only what is refused.

Stage 4. (1) A tool definition:

python
{"name": "get_invoice", "description": ( "Look up one invoice by " "its number. Use it before " "judging any dispute about " "that invoice. Returns the " "amount, dates and status."), "input_schema": { "type": "object", "properties": { "invoice_number": { "type": "string"}}, "required": [ "invoice_number"]}}

The description says what the tool does and when to use it, and the schema names one required input, so Claude has no guessing to do.

(2) A dispatcher for the six memory commands:

python
def handle(inp, store): cmd = inp["command"] p = inp.get("path") try: if cmd == "view": return store.view(p) if cmd == "create": store.write( p, inp["file_text"]) return "Created " + p if cmd == "str_replace": return store.replace( p, inp["old_str"], inp["new_str"]) if cmd == "insert": return store.insert( p, inp["insert_line"], inp["insert_text"]) if cmd == "delete": return store.delete(p) if cmd == "rename": return store.rename( inp["old_path"], inp["new_path"]) return ("Error: unknown " "command " + cmd) except Exception as e: return f"Error: {e}"

The memory tool is client-side: Claude requests each operation, and this function carries it out. Whatever it returns becomes the tool_result content. A failure, such as an unknown command, a missing key or a store method that raises because old_str is not in the file, comes back as a message Claude can read and correct its next call from, not as an exception that stops the loop.

(3) (a) Parallelization by voting: the same assessment runs several times, and a vote threshold decides the outcome. (b) Parallelization by sectioning: the refund check splits into independent subtasks that run at the same time, and their findings are combined. (c) Routing: each email is classified by language and directed to the prompt specialized for it.

Sources

Study Guide1,146 words

Agents and Workflows — study note

Read full article

Agents and Workflows — study note

This is Domain 1 of the Claude Certified Developer – Foundations exam: approximately 14.7% of scored items, according to the guide. The domain has three skills: agent architecture, agent construction with Claude, and agent patterns and frameworks. This note summarizes what each topic teaches and the Anthropic pages behind it.

Architecture: who decides the next step

Anthropic calls both kinds of system “agentic”, and separates them by who holds control.

DesignWho decides the next stepFits
WorkflowYour code, along predefined pathsWell-defined tasks where predictability matters
Orchestrator-workersA central model, per inputTasks whose subtasks cannot be predicted
AgentThe model, from what it observesOpen-ended problems with no fixed path
  • Start simple. Find the simplest solution, and add complexity only when it demonstrably improves outcomes. Often, a single model call with retrieval and in-context examples is enough.
  • Autonomy has a cost. Agents mean higher costs and the potential for compounding errors. Add stopping conditions, such as a maximum number of iterations, and test extensively in sandboxed environments.
  • Keep the agent grounded. It should judge progress from ground truth in the environment at each step, such as tool results or code execution. It can pause for human feedback at checkpoints or when it hits a blocker.
  • Subagents keep context clean. A subagent does a side task in its own context window, with its own tools and permissions, and returns only a summary.
    • Claude delegates from the subagent’s description, so write a clear one. A phrase such as “use proactively” encourages delegation.
    • If the tools field is omitted, a subagent inherits every available tool.
    • An optional memory field gives a subagent a persistent directory that survives across conversations.
    • Subagents work within a single session. For many independent parallel sessions, the subagents page points to background agents.

Construction: surfaces, the loop, hosting, hooks

SurfaceWho runs the agent loop
Agent SDKYour process, with Claude Code’s tools, permissions and sessions
Client SDKYour code; you write the loop, or use the beta tool runner
Managed Agents (beta)Anthropic’s hosted harness, in a managed or self-hosted sandbox

The Agent SDK in practice.

  • It loads skills, commands and memory from .claude/, the same as Claude Code.
  • It authenticates with API keys. Third parties may not offer claude.ai login without Anthropic’s approval.
  • From another language, run the CLI as a subprocess with -p and JSON output.

The loop. Call the API. While stop_reason is "tool_use", run every requested tool and send all the results back in one user message, each tool_result carrying the matching tool_use_id. Return a failing tool’s error with is_error: true instead of crashing. Claude can then retry with corrected input, ask for clarification, or explain the limitation.

Limits.

  • In the Agent SDK, each full cycle is one turn.
  • Cap turns, and set a budget: the agent-loop page calls a budget a good default for production agents. A capped run ends with an error_max_turns or error_max_budget_usd result subtype.
  • Read-only tools can run concurrently. Custom tools default to sequential until you set readOnlyHint.

Hosting an SDK agent. It is a long-lived process tied to local state, and one session maps to one subprocess. On-disk state does not survive a restart, scale-down or move to another node. Persist transcripts with a SessionStore adapter; give memory files and working-directory artifacts their own storage.

Session patternBest for
EphemeralOne-off tasks: a container per task, destroyed when it completes
HybridMany interactions with idle time between them
Long-runningAutonomous action, serving content, high-volume streams
Multi-agent containerAgents that must collaborate closely in a shared environment

Managed Agents (beta) is stateful by design: history, sandbox state and outputs are kept server-side. You keep control of that data: you can delete sessions, and separately delete any files you uploaded, at any time through the API. Because it is stateful, the overview currently notes it is not eligible for Zero Data Retention or HIPAA BAA coverage; check current eligibility before using it for regulated data.

Hooks.

  • Hooks are callbacks that run your code on agent events.
  • A matcher limits which tools a hook runs for. A hook without a matcher runs for every event of its type.
  • A PreToolUse callback can return allow, deny, ask or defer; ask shows the call to the user for approval. It can also modify the input or add context.
  • Async hook outputs let the agent continue without waiting, but they can’t block, modify or inject context, so use them only for side effects such as logging.
  • When hooks disagree, the most restrictive result applies, so a single deny blocks the call.

Patterns and frameworks

  • The tool-use contract. Claude emits a structured request. Your code, or Anthropic’s servers for server tools, runs it, and the result flows back. Claude sees only the schema and the result. If you are parsing a decision out of prose with a regex, make it a tool call.
  • Memory. The memory tool is client-side. Claude reads and writes files under /memories, and your application maps that prefix onto storage it controls. Restrict every operation to /memories.
PatternUse when
Prompt chainingThe task splits into fixed steps; add gates between them
RoutingCategories are distinct and classification is accurate
ParallelizationIndependent sections, or several attempts voting for confidence
Evaluator-optimizerClear criteria, and refinement adds measurable value

Routing can also cut cost, by sending easy questions to smaller models. Parallel sectioning suits guardrails: one call screens a query while another answers it.

Frameworks. Anthropic’s article notes that much of the tooling landscape described in the post has changed since December 2024; its advice is about how to work, not which tool to pick.

  • Frameworks make it easy to start by simplifying model calls, tool parsing and chaining.
  • They can also hide the underlying prompts and responses, and make it tempting to add complexity.
  • Start with the API directly, and if you adopt a framework, understand the code underneath.

Sources

Hands-on Lab2,715 words

Applications and Integration — practice exercise

Read full article

Applications and Integration — practice exercise

Difficulty: intermediate · Estimated duration: 75–95 minutes

You are the engineer responsible for ShelfScan, an invented internal service for a chain of hardware stores. Staff photograph shelves with a phone, and ShelfScan answers their questions about stock in a chat. You will set its requirements, plan for the model life cycle, write its request code, choose how its heavier jobs run, and then harden it: client and rate-limit handling, output and sessions, and the configuration that ships with it. Every stage can be completed without an API key and without spending anything: you read the material given here, write code or a decision, and check it against the reference solution. Stage 3 has an optional step: if you have your own API key you may run your function, which uses your own credit.

What you need: a text editor and a notes file. Python is used for the code; any language is fine for the reasoning.

This exercise covers all of Domain 2: Stages 1–4 cover requirements, life cycle and API mechanics; Stages 5–7 cover engineering foundations, application design and configuration.

StageWhat you practiseMinutes
1Criteria, evaluations and a starting guide14
2A life-cycle plan12
3The request and the stream12
4Images, thinking, batches, platform10
5Client, limits and delivery14
6Instructions, output, sessions, plugins14
7Instruction files, settings, pins10

Stage 1 — Requirements into criteria

Skills: CCDVF-U2.T1.LO1.S1, CCDVF-U2.T1.LO1.S2, CCDVF-U2.T1.LO1.S3 Minutes: 14

The brief from the operations director reads: “ShelfScan should count stock accurately and talk to staff politely.” The team wants to start writing the prompt today.

  1. Name the two things the team should produce before the first prompt, in order.
  2. Rewrite the brief as two success criteria: one quantitative criterion for counting and one criterion for politeness. The politeness criterion may use a qualitative scale; say what must be true of how that scale is applied.
  3. Which kind of evaluation fits a subjective judgement such as politeness, and what does it use to judge?
  4. ShelfScan is a chat. Name the edge case the criteria guide lists specifically for chat applications, and write one test case of that kind.
  5. How should you structure the evaluation's questions where possible, and why?
  6. Where can the team find a production guide to adapt for the staff chat, and which one fits best?

Stage 2 — Plan for the life cycle

Skills: CCDVF-U2.T1.LO2.S1, CCDVF-U2.T1.LO2.S2, CCDVF-U2.T1.LO2.S3 Minutes: 12

  1. How will the team hear that ShelfScan's model is scheduled for retirement?
  2. Once the model is deprecated, the director suggests waiting until the week before the retirement date to move. Give the documented reason to move sooner.
  3. ShelfScan calls the API with an SDK. Who sends the anthropic-version header, and what promise does Anthropic make about integrations that use the API as documented?
  4. After the move, ShelfScan's replies are noticeably shorter and more direct. Is something broken, and what should the team review?
  5. ShelfScan's old model had no effort parameter. What latency risk comes with leaving effort unset on the new one, and which output limit should the team revisit for requests that ran without thinking?

Stage 3 — Write the request

Skills: CCDVF-U2.T2.LO3.S1, CCDVF-U2.T2.LO3.S2 Minutes: 12

ShelfScan keeps a list history per chat. Standing rules for the assistant are in RULES.

  1. Write ask(history, text) in Python. It adds the staff member's message, calls the Messages API with the rules as standing instructions, adds Claude's reply to history so the next call has it, and returns the reply's text. Select the text block by its type.
  2. Some ShelfScan answers are long. Why might the SDK insist on streaming for requests with a large max_tokens, and how can you still get one complete Message back?
  3. The Python SDK offers two styles of stream. Which, and which suits a web server handling many chats at once?
  4. A support engineer asks which workspace ShelfScan's key resolved to for a failed request. Where can you read that from the response?
  5. (Optional, uses your own API key and credit.) Run ask twice in one chat and check that the second answer can use the first.

Stage 4 — Photos, thinking, batches and platform

Skills: CCDVF-U2.T2.LO3.S3, CCDVF-U2.T2.LO3.S4 Minutes: 10

  1. A staff member uploads an animated GIF of a shelf being restocked and asks what changed. What will Claude see?
  2. The director wants ShelfScan to report exact counts of small screws in a bin from one photo. What should the team tell them, and what does that mean for the design?
  3. How does Claude measure an image for token cost, and what does that suggest about very large photos?
  4. A nightly job re-checks every store's photos. It uses a manual thinking budget on its model. How must that budget relate to max_tokens?
  5. The nightly job moves to the Message Batches API. Can it keep the speed setting it used for fast mode? And how do you delete a batch that is still running?
  6. One region runs ShelfScan on Google Cloud and wants request-response logging. Does turning it on give Google or Anthropic access to the content?

Stage 5 — Client, limits and delivery

Skills: CCDVF-U2.T3.LO4.S1, CCDVF-U2.T3.LO4.S2, CCDVF-U2.T3.LO4.S3, CCDVF-U2.T3.LO4.S4 Minutes: 14

ShelfScan's nightly re-check now runs thousands of requests, and the team wants it to pace itself rather than collect 429s.

  1. Write call(client, **kw) in Python. It makes one Messages request through the raw-response interface, reads the remaining input tokens from the response headers, sleeps for PAUSE seconds when fewer than LOW remain, and returns the parsed message.
  2. The same script logs each response. Which helper turns a response into a plain dictionary? And after an SDK upgrade, a documented helper is missing: what should you check first?
  3. The async version of call uses the same raw-response interface. What changes about reading the parsed body?
  4. How precise is the remaining-tokens header, and in what format does the reset header give its time?
  5. A long report request returns a 504 timeout_error. What does the errors page suggest?
  6. ShelfScan's review workflow in GitHub Actions runs on every pull request. Name three documented ways to keep its runs from costing more than planned.
  7. Before merging a change Claude made, what should a reviewer ask Claude for, and how can Claude help with missing tests?

Stage 6 — Instructions, output, sessions and plugins

Skills: CCDVF-U2.T4.LO5.S1, CCDVF-U2.T4.LO5.S2, CCDVF-U2.T4.LO5.S3, CCDVF-U2.T4.LO5.S4 Minutes: 14

  1. Store managers use Claude Code in the terminal, in VS Code and in the desktop app. How many instruction files does the team need for the three surfaces, and why?
  2. Managers also want a weekly slide deck of stock trends. What does the prompting guide suggest adding to that request?
  3. Write a JSON schema for one shelf check: sku (string), count (integer), confidence (one of high or low) and note (string, default empty). Keep to supported features.
  4. Where in the response does the JSON arrive, and what does the Python SDK's messages.parse() do with output_format?
  5. ShelfScan's restock agent (TypeScript, Agent SDK) keeps a chat going across several messages. How does it continue the session without tracking IDs? And when the agent later moves from a laptop to a CI worker, can the worker resume from the session ID alone?
  6. A manager wants a plugin from Anthropic's official marketplace. Where can they see its context cost first, and is Claude Marketplace, the website, something they add with /plugin?

Stage 7 — Instruction files, settings and pins

Skills: CCDVF-U2.T5.LO6.S1, CCDVF-U2.T5.LO6.S2, CCDVF-U2.T5.LO6.S3, CCDVF-U2.T5.LO6.S4 Minutes: 10

  1. ShelfScan's repository already has a CLAUDE.md. What does /init do now, and what happens if two instruction files disagree?
  2. A developer answers “Yes, and don't ask again” to a Bash prompt in the CLI. Which file records it? What does the $schema line in a settings file add?
  3. A developer wants to try another model for one session without changing their default. How, in the /model picker? And does the Google Cloud deployment need a different ID format from the Claude API?
  4. ShelfScan's internal plugin depends on stock-tools at ^2.0, and the maintainers publish 2.1.0-beta.1. Will users get it? How can the maintainers preview a release tag without creating it?

Acceptance checks

  • Stage 1 puts criteria and evaluations before the prompt, writes one measurable and one consistently applied qualitative criterion, names an LLM-based Likert scale, the chat edge case, automated grading and the customer support agent guide.
  • Stage 2 cites email and documentation notices, the reliability of deprecated models, the SDK-sent version header and the as-documented promise, a prompt review for the more concise style, and the effort and max_tokens reviews.
  • Stage 3's ask sends the whole history with the rules in system, appends both turns, selects text by type, and the answers cover streaming for large outputs, both stream styles and the workspace response header.
  • Stage 4 answers each question with a mechanism, never with a price or a model name.
  • Stage 5's call reads the remaining-tokens header through the raw-response interface and returns the parsed message; the answers cover to_dict(), the installed version, awaiting on the async client, header precision and format, streaming after a 504, three cost controls, and risks and edge cases.
  • Stage 6 keeps one instruction file for every surface, asks for design elements in the deck prompt, writes a schema with only supported features, reads the text block, and covers continue: true, local session files, the context cost estimate and Claude Marketplace.
  • Stage 7 covers /init improvements, conflicting files, the local file, $schema, the picker's s key, the Google Cloud format, pre-release ranges and --dry-run.

Reference solution

Stage 1. (1) Success criteria first, then evaluations that measure performance against them. The prompt comes after both. (2) For example: “Counts match a staff recount on at least 95% of a held-out set of shelf photos” and “Replies score 4 or higher on a 1–5 politeness scale on at least 90% of test chats.” The qualitative scale is valuable only if it is applied consistently, along with quantitative measures. (3) An LLM-based Likert scale: a psychometric scale that uses an LLM to judge subjective attitudes or perceptions. (4) Poor, harmful or irrelevant user input. For example, a staff member asks ShelfScan to write a complaint about a colleague. (5) Structure them for automated grading, such as multiple choice, string match, code-graded or LLM-graded, so the set can grow without hand grading. (6) Anthropic's use-case guides, which are in-depth production guides for common use cases. The customer support agent guide fits: it covers context-aware chatbots for support interactions.

Stage 2. (1) Impacted customers are always notified by email and in the documentation. (2) Deprecated models are likely to be less reliable than active ones; moving to an active model keeps the highest level of support and reliability. (3) The SDK sends it automatically. If you use the API as documented, Anthropic generally will not break your usage. (4) Not necessarily: newer models have a more concise, direct communication style, so review the prompts against the prompting best practices. (5) Higher latency at the default effort level, so consider adjusting effort as you migrate; and review max_tokens for workloads that previously ran without thinking.

Stage 3. (1) A good answer:

python
import anthropic client = anthropic.Anthropic() def ask(history, text): history.append( {"role": "user", "content": text}) r = client.messages.create( model=MODEL, max_tokens=1024, system=RULES, messages=history) history.append( {"role": "assistant", "content": r.content}) return next( b.text for b in r.content if b.type == "text")

The API keeps nothing between calls, so history must carry both turns. Selecting by type survives replies that begin with thinking blocks. (2) The SDKs require streaming for large max_tokens to avoid HTTP timeouts. They can stream internally and still return the complete Message. (3) Sync and async streams; async suits a server handling many chats at once. (4) The anthropic-workspace-id response header, which names the workspace the key resolved to. (5) The second answer should refer to the first.

Stage 4. (1) Only the first frame; animations are unsupported. (2) Counts of many small objects are approximate, not precise. Design for approximate counts, for example a range or a prompt to recount by hand. (3) In patches rather than pixels; each patch is a visual token, so very large photos cost more tokens. (4) The budget must leave room for the final response, because thinking tokens count toward the turn's max_tokens. (5) No: fast mode tunes synchronous latency, which does not apply to batch processing, so drop speed. To delete a running batch, cancel it first. (6) No. Turning on the logging service gives neither Google nor Anthropic access to the content.

Stage 5. (1) A good answer:

python
import time def call(client, **kw): raw = (client.messages .with_raw_response .create(**kw)) left = int(raw.headers.get( "anthropic-ratelimit-" "input-tokens-remaining", "0")) if left < LOW: time.sleep(PAUSE) return raw.parse()

House guidance: pace before the limit instead of waiting for a 429. (2) to_dict() (or to_json()) on the response. If a helper is missing after an upgrade, the environment is probably still using an older version; print anthropic.__version__. (3) On the async client the raw response is an AsyncAPIResponse, so .parse() must be awaited. (4) Rounded to the nearest thousand; the reset time is in RFC 3339 format. (5) The request timed out while processing; consider the streaming Messages API for long-running requests. (6) Any three of: --max-turns in claude_args, workflow-level timeouts, GitHub's concurrency controls, specific requests, issue templates, a concise CLAUDE.md. (7) Ask Claude to highlight potential risks or considerations, and to identify edge cases you might have missed; its tests follow the style of your existing test files.

Stage 6. (1) One: the agentic loop, tools and capabilities are the same on every surface, so one project CLAUDE.md serves them all. (2) Thoughtful design elements, visual hierarchy, and engaging animations where appropriate. (3) A good answer:

json
{ "type": "object", "properties": { "sku": {"type": "string"}, "count": {"type": "integer"}, "confidence": { "type": "string", "enum": ["high", "low"]}, "note": { "type": "string", "default": ""} }, "required": ["sku", "count", "confidence", "note"], "additionalProperties": false }

(4) In the response's text content block. The Python SDK accepts output_format as a convenience and translates it to output_config.format. (5) Pass continue: true on each later query() call. No: session files are local to the machine that created them, so the ID alone is not enough on the worker. (6) In the Marketplaces tab of /plugin, where official-marketplace plugins show a context cost estimate. No: a plugin marketplace isn't Claude Marketplace.

Stage 7. (1) /init suggests improvements rather than overwriting the file. If two files give different guidance, Claude may pick either, so remove the conflict. (2) Only the local settings file. The $schema line gives autocomplete and inline validation in editors that support JSON schema. (3) Press s in the /model picker to switch without saving a default. No: on Google Cloud the format matches the Claude API. (4) No: a range doesn't match pre-release versions unless it opts in with a pre-release suffix. Pass --dry-run to claude plugin tag.

Sources

Study Guide2,290 words

Applications and Integration — study note

Read full article

Applications and Integration — study note

This note covers Domain 2 of the Claude Certified Developer – Foundations exam, the largest domain. It has six skills, all covered here:

  • understanding requirements;
  • the systems life cycle;
  • Claude API mechanics;
  • software engineering foundations;
  • Claude application design;
  • configuration management.

It summarizes what each topic teaches and the Anthropic pages behind it. On purpose, it teaches mechanisms rather than numbers. Model names, prices, size limits, retirement dates and rate limits change often, so look them up on the relevant pages when you need them. Where it gives general engineering advice beyond those pages, it says so.

From requirement to success criteria

Building a successful application starts with clearly defining success criteria and then designing evaluations to measure performance against them. Anthropic calls this cycle central to prompt engineering. Its prompt engineering overview assumes you already have criteria and a way to test against them; if not, establish those first.

Requirement sounds likeCriterion class
Same answer to a rephrased questionConsistency
Builds on what was said earlierContext utilization
Right language for the audienceTone and style
On the question, easy to followRelevance and coherence
  • Define the words. A criterion such as “90% of errors cause inconvenience, not egregious error” needs definitions of inconvenience and egregious before anyone can grade against it.
  • Qualitative is allowed. Qualitative measures can be valuable when applied consistently alongside quantitative ones.
  • Stuck? Brainstorm success criteria with Claude, with the criteria guide as guidance.

Evaluations that fit the criterion

CriterionMethodTest cases
Summary qualityROUGE-LReference summaries
ConsistencyCosine similarityGroups of paraphrases
Context useLLM ordinal scaleMulti-turn chats
Subjective toneLLM Likert scaleTarget-tone inquiries
  • Mirror the task. Evaluations should mirror the real task distribution, edge cases included: irrelevant or nonexistent input, overly long input and, for chat, poor, harmful or irrelevant user input.
  • Automate grading where you can, and get Claude to help generate more cases from a baseline set.

Production guides

Anthropic publishes production guides for common use cases: ticket routing, customer support agent, content moderation, legal summarization and a commerce agent blueprint. House advice: treat a guide as a starting set of requirements and adapt it to yours.

The model life cycle

StatusMeaning
ActiveFully supported and recommended
LegacyNo more updates; may be deprecated
DeprecatedWorks, not recommended; replacement and retirement date assigned
RetiredNo longer available; requests fail
  • Plan the move. Migrate all usage before the retirement date, and test the replacement well before it. Deprecated models are likely to be less reliable than active ones.
  • Find the usage. Export usage from the Console's Usage page for a CSV broken down by API key and model.
  • Parameters deprecate too. How the API treats a deprecated parameter depends on the model. Some SDKs remove parameters outright; most keep them in their types, so a clean type-check proves little.
  • Partner platforms set their own lifecycle dates, which can differ from the Claude API schedule.

API versions

Every request sends an anthropic-version header; the client SDKs send it for you. Anthropic recommends the latest version.

Preserved within a versionMay still change
Existing input parametersNew optional inputs
Existing output parametersNew output values
Error-type conditions
New enum variants

If you use the API as documented, Anthropic generally will not break your usage. Write code that tolerates additions.

Migrating to a newer model

Each model ID is a pinned version; updates ship under new IDs, and the guarantee covers IDs, not convenience aliases. Read the migration guide for the model you move to. Recurring patterns:

Old habitOn newer models
Prefill to continue a replyContinuation in a user message
Prefill to skip a preambleSystem-prompt instruction
Prefilled persona remindersReminders in the user turn
No thinking by defaultDisable thinking explicitly
Hand-parsed tool JSONA standard JSON parser

Also review prompts, since newer models are more concise and direct, and review max_tokens for workloads that ran without thinking.

Messages, data and streams

  • Stateless. Send the full conversation every time. Earlier assistant turns may be synthetic. Standing instructions go in the top-level system field.
  • Data access. The Models API lists models; the Files API stores a file once for reuse; newer lists page with page and limit, while batch and model lists use after_id and before_id. SDK auto-pagination only goes forward. A request over the size limit returns 413.
  • Streams. message_start carries an empty content; replies arrive in content-block events, reasoning in thinking_delta events, top-level changes in message_delta events.
  • Tool arguments arrive one complete key and value at a time; fine-grained streaming can be enabled per tool.
  • Recovery. Capture what arrived, send a continuation (a user message on newer models), resume. Prefer the SDK's accumulation.

Images and thinking

Image sourceUse it when
base64Any platform; one-off images
URLImage hosted online
file_idReused images, long sessions
  • On partner-operated platforms, only base64 sources are currently available.
  • Label several images, put them before the text, and send anything not in the pixels (such as metadata) as text.
  • Claude cannot reliably detect AI-generated images; counts are approximate; animations use the first frame.
  • Heavy compression can hurt accuracy; inspect the images you actually send.
  • Moving from a manual thinking budget to adaptive thinking is a behavioural change: adaptive thinking decides whether to think. The response usage reports thinking tokens; when streaming, on the final message_delta only.

Realtime, batch and cloud platforms

  • Batch what no one is waiting on. Requests run independently; streaming and fast mode are refused; params are validated at the end, so dry-run one request first; canceled batches keep partial results.
  • Google Cloud puts the model in the URL and the version in the body. Global endpoints suit flexible residency, multi-region a broad geography, regional a single region or provisioned throughput.
  • Amazon Bedrock runs inside the AWS security boundary; a service role is the recommended long-lived authentication.

Requests, SDKs and clients

HeaderNeeded
API versionAlways
Content typeAlways
Workspace IDSome keys
  • Authentication. Authorization: Bearer carries an API key or a short-lived token; x-api-key is a legacy fallback that still works. A federation token picks its workspace at exchange, so it never takes the workspace header. An SDK sends authentication, version and content type for you.
  • SDK caveats. The extra_ request options override documented parameters of the same name, so use them only with trusted input. A field returned as null and a field never returned both read as None; model_fields_set tells them apart.
  • Clients. Use the async client for concurrent services, and the aiohttp backend for better async performance. Timeouts raise APITimeoutError and are retried twice by default. Override retries or timeouts for one call with with_options. Close clients you create; pass custom HTTP clients as DefaultHttpxClient.
  • Shell scripts have an official tool: the ant CLI.

Rate limits and errors

Usage fieldCounts?
input_tokensYes
Cache writesYes
Cache readsNot usually
  • Token bucket. Capacity refills continuously, not at fixed intervals, and short bursts can exceed a limit. After a rate-limit 429, wait as long as retry-after says.
  • Output limits count tokens as they are generated; max_tokens does not count against them.
  • Headers show the limit, what remains and when it resets; the tokens headers show the most restrictive limit in effect, such as a workspace's.
  • Workspaces can have lower limits, except the default workspace, and organization limits always apply. Limits are per model, and are ceilings, not minimums.
  • Errors. Retry a 500 with backoff, then contact support with the request ID. A 504 suggests streaming. Error type values can grow, and a stream can fail after a 200.

Claude Code in the delivery pipeline

  • Refactor in small, testable increments that keep behaviour; tests follow your existing patterns; review each pull request and ask Claude for risks. claude --from-pr finds sessions linked to a pull request.
  • GitHub Action credentials. Share a Console API key, not a personal OAuth token; better, use workload identity federation with id-token: write. Deleting a secret leaves its key valid; delete the key in the Console too.
  • Scope each run. A plain-text prompt has no shell or GitHub access until you grant tools. Keep the claude_args line that starts the inline-comment server. Cap turns, set workflow timeouts and limit concurrency. The official GitHub App's permissions are all-or-nothing; a custom app can be narrower.

Instructions that reach Claude

  • Role and scope. A one-sentence role in the system prompt focuses behaviour; ask explicitly for fuller work; state identity and exact model strings when they matter; give design guidance for frontends.
  • Tools. Independent calls run in parallel; keep dependent calls sequential and forbid guessed parameters; ask for sequential steps if parallel work strains a system.
  • Safety. Consider reversibility: act on local, reversible steps, ask before hard-to-reverse or shared ones, and never take a destructive shortcut such as skipping a safety check.
  • Claude Code. Its system prompt isn't published; standing instructions go in CLAUDE.md or --append-system-prompt. The agentic loop is the same on every surface, so one set of instructions serves them all.

Output schemas

Schema featureSupported?
Local $refYes
External $refNo
RecursionNo
  • JSON outputs arrive in the text content block. Unsupported features return a 400: allOf with $ref, enums of objects, lookaround and backreference patterns, array bounds beyond minItems of 0 or 1.
  • They combine with batches and token counting, not with assistant prefilling.

Sessions

NeedUse
One-process chatSDK client
After a restartContinue
Many usersResume by ID
  • Read the session ID from the result message, present on success or error; a process failure yields none. Resume after a turn-limit error with a higher limit.
  • A fork copies history, not files: its edits land in the shared directory. persistSession: false keeps a TypeScript session in memory. Permission prompts inside one query() call are handled in the loop.

Plugins

  • Use a plugin to package several skills, agents, hooks or servers; a single skill works on its own. A marketplace is a catalog; plugins load at startup or on reload.
  • Scopes. User: you, everywhere. Project: the whole repository through committed settings, though each person installs it. Local: you, one repository. Cloud sessions skip local plugins.
  • Enabled plugins cost context every turn; disable unused ones. Any plugin runs with your privileges; official names count only from Anthropic's repositories. Managed settings can allowlist marketplaces and force-install plugins.

Instruction files and settings

  • AGENTS.md. Read alone when there is no CLAUDE.md; with both, only CLAUDE.md is read unless it imports AGENTS.md.
  • CLAUDE.md hygiene. Files are concatenated, so conflicts may resolve either way. HTML comments are stripped. A local file lives in one worktree; share personal notes by importing from your home directory. Exclusions skip other teams' files, never managed ones.
  • Settings. Managed settings win, then --settings for one session, then the local file, then the shared project file, then user settings. Files are strict JSON; a bad entry is skipped; most edits apply live, some keys only at start.

Pinning models and plugin versions

  • Model IDs are pinned per platform: Bedrock has its own format, Google Cloud matches the API, and each ID has its own retirement schedule. Pin a repository's model in shared settings; the startup header names the file that set it. Switching models mid-session re-reads the conversation uncached.
  • Plugin dependencies track the latest release unless you declare a tested semantic-version range. Ranges resolve against git tags, so maintainers must tag releases. Cross-marketplace dependencies need the root marketplace's allowlist, and conflicting ranges fail the install.

Sources

Study Guide1,066 words

CCDV-F exam map — domains, weights and how to use this hive

Read full article

CCDV-F exam map

This note maps the Claude Certified Developer – Foundations exam onto this hive. For each domain it shows how much of the exam it carries, where it is taught here, and how the mock papers sample it. Every exam fact below comes from the official exam guide, version 1.0 (effective July 2026). Where this hive makes its own choice, such as how many seats a domain gets on a mock paper, the note says so.

Who can take the exam

Anthropic's certification FAQ says: "Certification is currently available only to organizations in the Claude Partner Network." Registration also needs a partner email address on a recognized company domain. Personal email addresses do not work. Check that your organization is in the network before you plan around a date.

The FAQ also lists Claude Certified Developer – Foundations among the exams that count towards Claude Partner Network eligibility, so passing it adds to your organization's standing in the network. The same FAQ lists Claude Certified Associate under the exams that do not count towards that eligibility. If partner-network standing is your goal, this developer credential counts and the Associate one does not.

Registration goes through the Anthropic Partner Academy, and Pearson VUE delivers the exam. The guide states that there are no mandatory prerequisites. It recommends one to five years of software engineering and at least six months of hands-on work with Claude or comparable LLM-based systems, and says that experience is recommended, not required.

The exam at a glance

WhatExam guide v1.0
Items53
Item formatMultiple-choice and multiple-response; each item states how many responses to select
Time limit120 minutes
DeliveryProctored, online or at a test center, per program policy
Passing scoreScaled score of 720 on a scale of 100–1,000
Validity12 months from the date the credential is awarded

The guide sets the fee and the retake rules, and both can change. Check them in the current guide rather than here.

The eight domains

The guide publishes each domain's weight to one decimal place and calls the weights "the approximate proportion of scored items drawn from each domain". The mock-paper seats are this hive's own allocation: the guide's weights rounded to whole points, applied to 53 items and rounded by largest remainder, so the seats always add up to 53.

DomainWeightSkillsSeats per mock (house)
1 Agents and Workflows14.7%38
2 Applications and Integration33.1%617
3 Claude Code3.1%12
4 Eval, Testing, and Debugging2.6%11
5 Model Selection and Optimization16.8%49
6 Prompt and Context Engineering11.0%36
7 Security and Safety8.1%44
8 Tools and MCPs10.6%36

Domain 2 carries a third of the exam on its own. Domains 3 and 4 together carry under six percent, so a mock gives them only three seats between them. Spend your time where the weight is, but do not skip the small domains: every skill appears on every paper.

Where each domain is taught

Every domain has:

  • a narrated lecture deck for each topic, with a companion quiz;
  • flashcards;
  • a study note;
  • a hands-on code exercise with a reference solution.
DomainTopics in this hive
1 AgentsAgent and workflow architecture; Building agents with Claude; Agent patterns and frameworks
2 ApplicationsRequirements and the systems life cycle; Claude API mechanics; Software engineering foundations; Designing Claude applications; Configuration management
3 Claude CodeOperating Claude Code
4 Eval and debuggingDebugging and error handling
5 ModelsLLM fundamentals; Technical fundamentals; Choosing models and managing cost
6 Prompts and contextContext engineering; Prompt engineering and output handling
7 SecuritySecuring Claude applications; Guardrails, hooks and secrets
8 Tools and MCPsImplementing tools; MCP servers and agentic customization

How to use the hive

  1. Learn a topic. Watch the deck, then take its quiz straight away. The quiz tests only what that deck teaches.
  2. Consolidate a domain. Read the domain's study note and review its flashcards.
  3. Apply it. Work through the domain's code exercise, then compare your work with the reference solution. Every exercise can be finished without spending on the API; any optional step that calls the API says so and uses your own key.
  4. Sit a mock paper once you have covered all eight domains. Take it timed, with no notes.
  5. Review by domain. Note which domains your missed questions come from. Go back to the decks for your weakest domain, then sit the next paper.

The four mock papers

  • Each paper has 53 questions, is timed at 120 minutes, and includes multiple-response items that state how many to select, as the real exam does.
  • The four papers share no questions. Every paper asks about all 25 skills, in the seat counts shown above.
  • The exam-bank questions are separate from the quiz questions. Scoring well on a paper means applying what you learned to new situations, not recognizing a quiz item you have seen before.
  • A mock result is the percentage you answered correctly. It is not the exam's scaled score. The guide reports a pass or fail on a scaled score of 100–1,000 with a cut score of 720. It also reports the percentage correct in each domain, and that per-domain breakdown is the part a mock can imitate.

Staying current

The Claude platform changes faster than the exam. This hive teaches decisions that last: when to cache, how to handle a stop reason, where a credential should live. It avoids prices, rate limits, context-window sizes and model version numbers. When a model, parameter or tool on the platform differs from a deck, trust the documentation page the deck cites, which is linked from the hive's sources.

Hands-on Lab1,111 words

Claude Code — practice exercise

Read full article

Claude Code — practice exercise

Difficulty: intermediate · Estimated duration: 45–60 minutes

You will write the commands, files and choices a developer makes to run Claude Code well on a small project, and check them against the reference solution. Everything is written down, so no API or subscription spend is needed. If you have Claude Code installed and want to see a step working, running it is optional and uses your own account.

The project: an invented order-tracking service in a git repository, with a test suite, a weekly release procedure and a CI pipeline.

StageWhat you practiseMinutes
1Steering the loop, managing context, resuming or forking12
2Memory, a skill and a read-only subagent14
3A headless CI step and its permission mode12
4Settings scopes and two workflows12

Stage 1 — Context and sessions

Skills: CCDVF-U3.T1.LO1.S1 Minutes: 12

For each situation, write the command or action you would take and one sentence saying why.

  1. Claude has started rewriting the whole order parser, but you only wanted the currency rounding fixed.
  2. You are deep into renaming the order statuses. The window is nearly full, and you want the summary to keep the list of files still to rename.
  3. An edit Claude made to the order model broke three tests, and you want the code and the conversation back to how they were before it.
  4. You want to see what is taking up the context window.
  5. Yesterday's session has a working fix. You want to try a riskier approach today without disturbing it.

Stage 2 — Memory, skills and subagents

Skills: CCDVF-U3.T1.LO1.S2 Minutes: 14

  1. The team pastes the same 60-line release checklist into chat every Friday. Write the path of a skill that holds it (skills live at .claude/skills/<name>/SKILL.md) and its first lines: frontmatter with a description, then the start of the body.
  2. Write the frontmatter for a project subagent, saved under .claude/agents/, that audits the repository for hard-coded secrets. It must be able to read and search files (Read, Grep, Glob) but never edit or write them.
  3. In two sentences, explain to a teammate why the checklist should not live in CLAUDE.md, and what auto memory is for instead.

Stage 3 — A headless CI step

Skills: CCDVF-U3.T1.LO1.S3 Minutes: 12

  1. Write a shell step for CI that asks Claude Code to run the test suite (npm test) and explain any failures. The step must:
    • stream its progress as newline-delimited events, so the CI log shows it live;
    • never wait for an approval (set the mode with --permission-mode);
    • pre-approve only the test command;
    • fail the pipeline if the run fails.
  2. Name the permission mode you would use in each place, with one reason:
    • (a) your laptop, while you review every change;
    • (b) the CI runner;
    • (c) a throwaway container with no network access, for an experiment.

Stage 4 — Settings and workflows

Skills: CCDVF-U3.T1.LO1.S4 Minutes: 12

  1. Put each setting in the right file:
    • (a) the team's shared permission rules;
    • (b) your preferred theme, in every project;
    • (c) a model override you want only in this project.
  2. Before a risky schema change, you want Claude to read the code and propose a plan without touching disk. Write the command that starts the session that way.
  3. You need to learn how orders are archived, but you do not want dozens of file reads in your main context. Write the one-line request you would send.

Acceptance checks

  • Every Stage 1 answer names a command or action and says why.
  • The skill sits under .claude/skills/ with a clear description (frontmatter fields are optional; description is the one that matters most), and the subagent restricts its tools with an allowlist.
  • The CI step uses -p, requests streamed output (stream-json with --verbose), uses a mode that refuses prompts, pre-approves only the test command, and fails on a non-zero exit.
  • Each setting is in a file whose scope matches who it should apply to.
  • The plan-mode command and the delegation request each match the workflow the stage asks for.

Reference solution

Stage 1

SituationAction
1Interrupt and redirect to the rounding fix only
2/compact with a focus on the files still to rename
3/rewind to the checkpoint before the edit
4/context
5Fork yesterday's session, which leaves the original unchanged

You can interrupt at any point to steer. Compaction keeps requests and key snippets but can lose early detail, so focus it. /rewind rolls code and conversation back to a checkpoint; checkpoints cover file changes only, not remote systems. Resuming would append to the original session; forking copies it into a new session ID.

Stage 2

Path: .claude/skills/release-checklist/SKILL.md

markdown
--- name: release-checklist description: >- Friday release steps for the order service --- 1. Confirm main is green...
markdown
--- name: secrets-audit description: >- Finds hard-coded secrets in the repository tools: Read, Grep, Glob ---

A skill's body loads only when it is used, while CLAUDE.md loads in every session. Auto memory is what Claude writes itself from your corrections, not a place for procedures. The tools field is an allowlist, so this subagent cannot edit or write files.

Stage 3

shell
claude -p "Run npm test and \ explain failures" \ --output-format \ stream-json --verbose \ --permission-mode \ dontAsk --allowedTools \ "Bash(npm test)" || exit 1
  • (a) acceptEdits: reads and edits go through while you review the changes.
  • (b) dontAsk: anything that would prompt is denied, which suits locked-down CI.
  • (c) bypassPermissions: only in an isolated container or VM like this one.

Stage 4

  • (a) .claude/settings.json, committed so everyone gets the same permission rules.
  • (b) ~/.claude/settings.json, which applies to you in every project.
  • (c) .claude/settings.local.json inside the project, which changes nothing for your teammates.

Start the session with claude --permission-mode plan: Claude reads files and proposes a plan, but makes no edits until you approve.

A good request: "Use a subagent to investigate how orders are archived and report a summary." The subagent reads files in its own context window and returns only the findings; its requests still count toward your usage limits.

Sources

Claude Code documentation: how Claude Code works, commands, memory, skills, subagents, headless mode, permission modes, settings and common workflows.

Study Guide797 words

Claude Code — study note

Read full article

Claude Code — study note

Domain 3 of the Claude Certified Developer – Foundations exam covers one skill, Claude Code Operation. It is a small share of the exam, and it rewards knowing when each mechanism applies, not memorizing flags. This note summarizes the topic deck and the Claude Code pages behind it.

The loop and the harness

Every task runs through gather context → take action → verify results. The phases blend, and you can interrupt at any point to steer. Claude Code is the layer around the model that provides the tools and manages the context the model sees; the documentation calls this layer the agentic harness.

Context and sessions

NeedMechanism
See what fills the window/context
Free space, same conversation/compact, optionally with a focus
Start over on a new task/clear (empty context)
Undo an edit and the turns after it/rewind to a checkpoint (file changes only)
A rule that must survive compactionCLAUDE.md

When the window fills, Claude Code clears older tool outputs first, then summarizes. To steer every compaction in a repository, add a Compact Instructions section to CLAUDE.md. Checkpoints cover file changes only; actions on remote systems such as databases, APIs or deployments can't be checkpointed. Your requests and key snippets are kept, but early detailed instructions may be lost. Each new session starts with a fresh context window. Resume appends to the same session. Fork copies the history into a new session and leaves the original unchanged.

Memory, skills and subagents

FeatureWritten byLoads
CLAUDE.mdYouEvery session
Auto memory (MEMORY.md)Claude, from your correctionsThe start of the file, up to a limit, every session
SkillYouDescription by default; body when used
SubagentYouRuns in its own context window

Where they live:

  • Skill: .claude/skills/<name>/ SKILL.md

  • Subagent: .claude/agents/

  • Both memory systems are context, not enforced configuration. Keep them short, because longer files reduce adherence.

  • Create a skill when you keep pasting the same procedure into chat.

  • Skill descriptions load by default, with two exceptions. With many skills, some descriptions are dropped to fit the listing's budget, least used first. And disable-model-invocation: true keeps a description out of context, so the skill runs only when you invoke it.

  • Auto-memory topic files are not loaded at startup; Claude reads them on demand.

  • Delegate exploration to a subagent to keep file reads out of your window. Its requests still count toward the same usage limits.

  • Restrict a subagent with the tools allowlist or the disallowedTools denylist.

Running without a person

  • claude -p runs a prompt non-interactively. It exits 0 on success and non-zero on failure, so scripts can branch on the exit status.
  • --output-format takes one of three values:
    • text, the default;
    • json, which gives the result, session ID and metadata, including an estimated cost (client-side, and it can differ from your bill);
    • stream-json, which gives newline-delimited events.
  • --allowedTools pre-approves the listed tools. Set a mode with --permission-mode, for example --permission-mode dontAsk.
  • Piped stdin has a size cap. Past it, the run exits with an error, so write large input to a file and reference its path.
ModeBest for
acceptEditsIterating on code you are reviewing
dontAskLocked-down CI: anything that would prompt is denied
bypassPermissionsIsolated containers or VMs only

Non-interactive runs start in the default (Manual) mode unless you pass a mode. Under dontAsk, actions that need no approval, such as file reads in the working directories, still run, and so do tools you pre-approved. Deny rules block in every mode. Allow rules have no effect in bypassPermissions, and a short list of actions is never auto-approved in any mode.

Settings and everyday workflows

Settings fileApplies to
UserYou, in every project on the machine
ProjectEveryone, once committed
LocalYou, in this project only

The files, in that order:

  • User: ~/.claude/settings.json

  • Project: .claude/settings.json

  • Local: .claude/settings.local.json

  • Worktrees give parallel sessions on separate branches (claude --worktree <name>). The repository needs at least one commit.

  • Plan mode proposes a plan and makes no edits until you approve.

  • Claude reads files fresh on each tool call, so it sees your manual edits.

  • Scheduled tasks run autonomously, so their prompts must say what success looks like and what to do with the results.

Hands-on Lab1,106 words

Debugging and Error Handling — practice exercise

Read full article

Debugging and Error Handling — practice exercise

Difficulty: intermediate · Estimated duration: 45–60 minutes

You will diagnose a set of failures from an invented application and write the code or configuration that handles each one. The failures are given as logs and response fragments, so no API spend is needed. You can write the code in Python or TypeScript. Running any of it against the live API is optional and uses your own key.

The application: an invoice-processing service that calls the Claude API, uses one server tool (web search) and two client tools (lookup_customer, post_ledger_entry), and runs a nightly agent job through the Agent SDK with tracing enabled.

StageWhat you practiseMinutes
1Classifying HTTP errors and choosing retry, fix or escalate12
2Handling the stop reasons that need code14
3Placing tool failures in the integration or the model12
4Reading a trace12

Stage 1 — Error triage

Skills: CCDVF-U4.T1.LO1.S1 Minutes: 12

For each response, write the class of error and what the client should do.

  1. 401 with authentication_error, on every call since a deploy.
  2. 529 with overloaded_error, on a few calls during the afternoon peak.
  3. 402 with billing_error, on every call since this morning.
  4. 409 with conflict_error, on an update to a resource another job changed first.

Then write the error-handling block for one call. It must catch the SDK's typed exception for status errors, branch on the status code, log the request ID, and leave transient retries to the SDK.

Stage 2 — Stop reasons

Skills: CCDVF-U4.T1.LO1.S2 Minutes: 14

  1. Write a function that takes a successful response and handles end_turn, max_tokens, pause_turn and refusal.
  2. Explain in one sentence why the function is called on responses whose HTTP status is 200.
  3. A response stops with tool_use for lookup_customer. Say how this differs from pause_turn, and what your code sends next.

Stage 3 — Integration or model?

Skills: CCDVF-U4.T1.LO1.S3 Minutes: 12

For each symptom, say whether the fault is in the integration or in what the model was given, and write the fix.

  1. A 400 says tool_use ids were found without tool_result blocks immediately after. Your code sent the two results in two separate messages, each after a line of text.
  2. Claude calls post_ledger_entry when you expected lookup_customer.
  3. Claude sometimes passes status: "archived" to lookup_customer, a value outside the tool's status enum of forty values.
  4. Your code puts the instruction "only look up active customers" inside the lookup_customer result, and Claude asks the user to confirm it.

Stage 4 — Read a trace

Skills: CCDVF-U4.T1.LO1.S4 Minutes: 12

The nightly job's trace for one turn shows:

  • the turn span: 94 seconds;
  • two Claude API call spans: 3 and 4 seconds;
  • one post_ledger_entry tool span: 86 seconds, whose permission-wait child is 85 seconds and whose execution child is 1 second.
  1. Where did the time go, and what would you change?
  2. The team added a hook that runs before each tool call. With detailed tracing on, which span shows the hook's own time?
  3. The job's spans appear as a separate trace from the application that started them. What connects them?

Acceptance checks

  • Every Stage 1 answer names the class and the action, and the code catches the SDK's status-error type, branches on the status code, and logs the request ID.
  • The stop-reason function branches on stop_reason, marks truncations as incomplete, resends on pause_turn, and routes refusals to stop_details and a fallback model.
  • Each Stage 3 fix is placed correctly: integration, description, schema, or where the instructions go.
  • The trace answers use what each span covers, not guesses.

Reference solution

Stage 1

ResponseClass and action
401The key: fix the credential, and do not retry
529Transient overload: retry with backoff
402Billing: fix the payment details, and do not retry
409Conflict: resolve it, then retry
python
from anthropic import APIStatusError try: msg = client.messages.create( **request) except APIStatusError as e: rid = e.response.headers.get( "request-id") log.error("status %s id %s", e.status_code, rid) code = e.status_code if code in (401, 402, 403): alert_owner() # fix access raise if code == 409: msg = reload_and_retry() else: raise

The SDK already retries connection errors, rate limits and 5xx errors with backoff, so the handler branches on status_code and re-raises rather than looping. 401, 402 and 403 alert the owner to fix the key, billing or access; 409 resolves the conflict before its retry. anthropic.APIStatusError and status_code are the names the stop-reasons page uses in Python, and the Python SDK page reads the ID with response.headers.get("request-id"). Error bodies carry the same ID as request_id. The helper names are placeholders.

Stage 2

python
NOTE = "\n[output incomplete]" def handle(msg, history): reason = msg.stop_reason if reason == "max_tokens": return text_of(msg) + NOTE if reason == "pause_turn": history.append({ "role": "assistant", "content": msg.content}) return None # loop resends if reason == "refusal": log.info("refusal: %s", msg.stop_details) return None # try fallback return text_of(msg) # end_turn

A refusal arrives as a normal HTTP 200, so the status alone would pass it on as an answer; the documented action is to read stop_details and retry on a fallback model. A tool_use stop waits on your own tool, so your code runs lookup_customer and sends a tool_result. pause_turn is a server-tool loop pausing, so you resend the response as-is.

Stage 3

  1. Integration: send one message whose content is both tool_result blocks, before any text.
  2. The model's input: sharpen the tool descriptions so they say when to use each tool.
  3. The model's input: shrink the enum, or add input_examples showing the valid choices.
  4. Integration: move the instruction out of the tool result and keep the result to data. The page names two places: a user turn after the tool_result block, or, on supported models, a mid-conversation system message.

Stage 4

  1. 85 of the 86 seconds were the permission wait, not the tool. Pre-approve the call for the nightly job, or run it in a mode suited to unattended work.
  2. The span that wraps each hook execution; it needs detailed beta tracing.
  3. Start the agent run while an application span is active. The SDK propagates W3C trace context, so the run appears inside the application's trace.

Sources

Claude API errors; handling stop reasons; troubleshooting tool use; the Python SDK; Agent SDK observability.

Study Guide733 words

Debugging and Error Handling — study note

Read full article

Debugging and Error Handling — study note

Domain 4 of the Claude Certified Developer – Foundations exam covers one skill, Debugging and Error Handling. The habit it rewards is placing a failure before fixing it:

  • a failed request (an error);
  • a successful response that stopped for a reason (a stop reason);
  • a fault in your integration;
  • a problem in what the model was given.

This note summarizes the topic deck and the pages behind it.

Errors: the status says who acts

StatusMeaningWhat to do
400Format or content of the request, or a spend limit you setFix the request, or review the limit
401 / 403The key, or its permissionFix the credential or access
409Conflict with a resource's current stateResolve it, then retry
413Request too largeShrink or split it
429Rate limit, or a spend capBack off; a spend-cap 429 keeps failing until access resumes
500 / 529Internal error / temporary overloadRetry with exponential backoff
  • Errors are JSON with an error object holding a type and a message, plus a request ID. Keep the request ID for support.
  • Catch the SDK's typed exceptions, most specific first. Never string-match messages.
  • The official SDKs already retry transient failures (connection errors, rate limits, 5xx) with backoff.
  • Streams can fail after a 200, and those errors arrive as events, not as an HTTP status.
  • If traffic jumps sharply, ramp it up gradually to avoid acceleration limits.

Stop reasons: a 200 still needs checking

stop_reasonWhat to do
end_turnUse the response
max_tokensRaise the limit, or continue the response
stop_sequenceRead stop_sequence to see which one fired
Context window fullTreat as truncated (see below)
pause_turnA server-tool loop paused: send the response back as-is
tool_useRun your tool and send a tool_result
refusalArrives as a normal HTTP 200: read stop_details and retry on a fallback model

The context-window row is the stop reason model_context_window_exceeded. Newer models return it; earlier models need a beta header.

When you truncate, append a notice so readers know the output is incomplete. An empty end_turn right after tool results usually means text was added after the tool_result blocks: stop adding it, and don't resend the empty response unchanged. When streaming, stop_reason arrives in the message_delta event.

Integration fault or model output?

SymptomWhere the fault livesFix
400: tool calls left without resultsYour integrationOne result per call, before any text
Tools called one per turnYour integrationAll results in one user message
Thinking blocks cannot be modifiedYour integrationSend the assistant message back unchanged
Wrong tool chosenThe descriptionsSay when to use each tool
Wrong parameter typesThe schemaStrict mode (supported subset) or input examples
Claude refuses to act on a tool result, or asks to confirm its instructionsWhere you put themMove them to a user turn after the result, or (supported models) a mid-conversation system message
String match on tool input breaksYour integrationParse the JSON: escaping differs by model version

In the API these are tool_use and tool_result blocks matched by id; the schema fixes are strict: true (if the schema is in the supported subset) and input_examples on the tool definition; the instruction fix is a user turn after the result.

Traces: turn, API call, tool call

With tracing enabled (in beta, so names may change), each step becomes a span:

  • the turn: one prompt to one response;
  • each API call: model, latency and token counts;
  • each tool call: with separate children for the permission wait and the execution.

Failed or aborted requests may omit token counts. Content is not recorded by default. Start an agent run while one of your application's spans is active, and it appears inside your application's trace.

Hands-on Lab1,534 words

Model Selection and Optimization — practice exercise

Read full article

Model Selection and Optimization — practice exercise

Difficulty: intermediate · Estimated duration: 45–60 minutes

You are the engineer responsible for TriageDesk, an invented internal service that reads incoming support tickets, tags each one and drafts a first reply. It is high-volume and latency-sensitive: thousands of tickets a day, and agents wait for each tag. You will review its design and code and write down the decision you would make at each step. Every stage can be completed without an API key and without spending anything: you read the material given here, write code or a decision, and check it against the reference solution. Stage 2 has an optional step: if you have your own API key you may run your rewritten client, which uses your own credit.

What you need: a text editor and a notes file. Python is used for the code; any language is fine for the reasoning.

StageWhat you practiseMinutes
1Reading cache usage, sizing effort, checking examples12
2Rewriting a hand-rolled client and streaming the reply14
3Choosing a transport and keeping a slow tool alive8
4A model and cost plan16

Stage 1 — Size the request

Skills: CCDVF-U5.T1.LO1.S1, CCDVF-U5.T1.LO1.S2, CCDVF-U5.T1.LO1.S3 Minutes: 12

TriageDesk logged this usage for one request (the numbers are invented for the exercise):

  • input_tokens: 3,100
  • cache_read_input_tokens: 41,000
  • cache_creation_input_tokens: 0
  • output_tokens: 2,400

The request carried a system prompt, 22 tool definitions, a ticket and its thread history. Thinking is enabled and the model supports the effort parameter. Most tickets are simple password resets.

  1. From the table, say how many input tokens were read from the cache and how many were written to it, and whether this request hit the cache.
  2. Choose one change for the simple tickets and justify it in one sentence: lower effort, or a larger thinking budget.
  3. The tagging prompt carries four example tags, all taken from password-reset tickets, and the team is unsure they are varied enough. What can you ask Claude to do with them?

Stage 2 — Harden the client

Skills: CCDVF-U5.T2.LO2.S1, CCDVF-U5.T2.LO2.S2 Minutes: 14

TriageDesk's current client, simplified:

python
import requests def draft(ticket): r = requests.post( API_URL, json=body(ticket), headers={"x-api-key": KEY}) return r.json()
  1. The client already sends the API key. Name three other things it leaves to chance that a client SDK handles for you.
  2. Rewrite draft to use the Python SDK. Its requests use a large max_tokens, and the agent UI does not show the reply until it is complete, so avoid holding a long idle connection. Return the reply's text and raise a clear error when there is none.
  3. Agents want a short view of Claude's reasoning while a draft streams, not the full chain of thought. Which thinking display setting gives that?
  4. The stream to the agent UI sometimes drops mid-reply, after a tool call and some text. Can you resume, and from where?
  5. (Optional, uses your own API key and credit.) Run your rewritten draft once against the API and compare the usage field with your Stage 1 reasoning.

Stage 3 — Pick the transport

Skills: CCDVF-U5.T2.LO2.S3 Minutes: 8

  1. TriageDesk's engineers keep a script that reads log files on their own machines, and want Claude Code to use it as an MCP server. Which transport fits, and why?
  2. The knowledge-search tool, on a remote MCP server connected to Claude Code, can work for minutes without sending anything, and its calls are aborted with an error before their overall time limit. What should the server do?

Stage 4 — Plan the model and the bill

Skills: CCDVF-U5.T3.LO3.S1, CCDVF-U5.T3.LO3.S2, CCDVF-U5.T3.LO3.S3, CCDVF-U5.T3.LO4.S1, CCDVF-U5.T3.LO4.S2, CCDVF-U5.T3.LO4.S3 Minutes: 16

Write a one-page plan that answers each question in two or three sentences:

  1. Starting model. TriageDesk is high-volume and latency-sensitive. Which starting approach fits, and why?
  2. Effort and cost. A tag can be checked against the ticket's final resolution. How do you set effort to keep cost down, and why might a pricier model still cost less per ticket?
  3. Upgrade. List three checks for moving TriageDesk to a newer model.
  4. Caching. The system prompt and tool definitions are cached. What must stay unchanged between requests for the cache to hold?
  5. Counting. Some tickets attach screenshots by URL. How do you count their tokens before sending?
  6. Tracking and batching. A nightly job re-tags last week's tickets. How do you run it, and how do you avoid double-counting cost in your Agent SDK worker?

Acceptance checks

  • Stage 1 reads the cache fields correctly, picks lower effort for simple tickets, and asks Claude to review the examples for relevance and diversity.
  • Stage 2 uses the SDK, streams internally and returns the complete message, selects content by type, picks the summarized display, and resumes from the most recent text block.
  • Stage 3 picks stdio for the local script and progress notifications for the slow tool, with the reason for each.
  • Stage 4 answers all six questions with a mechanism, never with a price or a model name.

Reference solution

Stage 1. (1) cache_read_input_tokens is the number of tokens retrieved from the cache for this request, so 41,000 input tokens came from the cache. cache_creation_input_tokens is the number written when creating a new entry, so nothing was written. The request hit the cache for its prefix, and only 3,100 input tokens were processed uncached. (2) Lower effort, since this model supports it. Effort applies to every output token, including tool calls, and it is the lever for trading thoroughness against tokens on a single model. A larger budget spends more. (3) Ask Claude to evaluate the examples for relevance and diversity, or to generate additional ones from the initial set, so the examples cover more than password resets.

Stage 2. (1) Any three of: the version and content-type headers; retries; timeouts; error handling; streaming helpers. (2) A good rewrite:

python
import anthropic # reads the key from the environment client = anthropic.Anthropic() def draft(ticket): with client.messages.stream( **body(ticket)) as s: msg = s.get_final_message() text = next( (b.text for b in msg.content if b.type == "text"), None) if text is None: raise ValueError( msg.stop_reason) return text

Some networks drop idle connections, so a large max_tokens without streaming can fail. The SDK streams internally and returns the complete message, which is what the UI needs. A reply can begin with thinking blocks, so the rewrite selects content blocks by type rather than by position. (3) display: "summarized", which streams a condensed summary of Claude's reasoning rather than the full chain of thought. (4) Yes, partly. Tool use and extended thinking blocks cannot be partially recovered; resume streaming from the most recent text block. (5) The response's usage field reports what the request consumed.

Stage 3. (1) Stdio. Stdio servers run as local processes on your machine and suit tools that need direct system access or custom scripts. (2) Send progress notifications while it works. A tool call that sends no response and no progress notification for the idle window aborts with an error instead of waiting for the wall-clock limit.

Stage 4. (1) Efficiency-first: start with a faster, more cost-effective model. The selection guide lists applications with tight latency requirements among those that approach is best for. (2) Run every ticket at a low effort setting and re-run only the failures at a higher one, which needs a checker that does not pass bad tags. Compare models on cost per completed task: a more capable model can finish with less work, with fewer turns and less backtracking. (3) Read the migration guide's breaking changes; select content blocks by type, since a reply can begin with thinking blocks; and re-baseline cost on your typical workload before production. (4) Keep the tool definitions unchanged, because editing them invalidates the entire cache. If you use a task budget, set it once on the first request, because changing it partway invalidates any cached prefix that contains it. Keep the effort level and thinking configuration fixed too, since changing them between requests invalidates the cache from that point onward. (5) Send the screenshots as base64 to the counting endpoint; it does not accept images given by URL. The count is an estimate. (6) Send it through the batch API, since no one is waiting on it, and keep the interactive path for live tickets. In the Agent SDK, deduplicate usage by message ID, because messages from parallel tool calls in one turn share an ID.

Sources

Ready to practice? Jump straight in — no sign-up needed.

Take practice tests, review flashcards, and read study notes right now.

Take a Practice Test

Claude Certified Developer - Foundations (CCDV-F) Practice Questions

Try 15 sample questions from a bank of 503. Answers and detailed explanations included.

Q1medium

A code-review bot must never edit files unless a reviewer explicitly asks it to, even when a comment is ambiguous. What should its system prompt set?

A.

A default to implement changes rather than only suggesting them, so reviewers save time on fixes

B.

A rule to ask a clarifying question before every single tool call it makes

C.

A default to inform and recommend when intent is ambiguous, editing only on explicit request

D.

A rule that it must never use any tool at all, stated in capitals for emphasis

Show answer & explanation

Correct Answer: C

A default to implement changes rather than only suggesting them, so reviewers save time on fixes — Incorrect: Defaulting to implementation is the opposite, proactive setting.

A rule to ask a clarifying question before every single tool call it makes — Incorrect: Questioning before every tool call also blocks the reads a review needs.

A default to inform and recommend when intent is ambiguous, editing only on explicit request — Correct: The conservative sample in the prompting guide: when intent is ambiguous, default to information, research and recommendations, and edit only when explicitly requested.

A rule that it must never use any tool at all, stated in capitals for emphasis — Incorrect: Forbidding all tools stops the bot reading the code it reviews.

Answer: C

Q2hard

A plugin’s own plugin.json declares a dependency from a marketplace that the root marketplace has not allowlisted. What does a user see on install?

A.

The install refuses with an allowlist message

B.

The dependency installs anyway, with a warning

C.

The plugin loads without the dependency and works

D.

The install completes, and the plugin then fails to load

Show answer & explanation

Correct Answer: D

The install refuses with an allowlist message — Incorrect: That refusal is for a dependency declared in the marketplace entry.

The dependency installs anyway, with a warning — Incorrect: The dependency is not installed.

The plugin loads without the dependency and works — Incorrect: The plugin fails to load.

The install completes, and the plugin then fails to load — Correct: When the dependency is declared in plugin.json, the install completes without it and the plugin then fails to load.

Answer: D

Q3medium

Which kind of product does the article give as a good example for orchestrator-workers?

A.

A coding product that makes complex changes to multiple files each time

B.

A form that always validates the same five fields

C.

A classifier that sends each message on to one of three specialized prompts

D.

A translation step followed by a fixed length check

Show answer & explanation

Correct Answer: A

A coding product that makes complex changes to multiple files each time — Correct: Coding products that make complex changes to multiple files each time are the article’s example.

A form that always validates the same five fields — Incorrect: A fixed set of checks is predefined, so it suits a workflow.

A classifier that sends each message on to one of three specialized prompts — Incorrect: That is routing.

A translation step followed by a fixed length check — Incorrect: That is prompt chaining.

Answer: A

Q4medium

Before each deploy command, an agent should see a reminder that a release freeze is on, but the command should still run. What can the PreToolUse hook callback return?

A.

A deny, since blocking is how a hook changes what the agent sees

B.

A new name for the deploy tool for the length of the freeze

C.

Context injected into the conversation, while allowing the call

D.

A log line, which the model reads before its next turn

Show answer & explanation

Correct Answer: C

A deny, since blocking is how a hook changes what the agent sees — Incorrect: Deny blocks the command, which the team does not want.

A new name for the deploy tool for the length of the freeze — Incorrect: Hooks do not rename tools.

Context injected into the conversation, while allowing the call — Correct: A callback can allow the operation, block it, modify the input, or inject context into the conversation.

A log line, which the model reads before its next turn — Incorrect: The model reads what the hook returns, not your logs.

Answer: C

Q5medium

A service uses with_streaming_response to read a response header before the body, but calls it without a with block, and open connections pile up. What does the documentation require?

A.

Calling .parse() first, which reads the body and always closes the stream

B.

Setting max_retries=0, since retries hold connections open

C.

Using it as a context manager, so the response is reliably closed

D.

Switching to the async client, which closes streamed responses by itself

Show answer & explanation

Correct Answer: C

Calling .parse() first, which reads the body and always closes the stream — Incorrect: Reading the body is not what guarantees the response is closed.

Setting max_retries=0, since retries hold connections open — Incorrect: Retries are unrelated to closing a streamed response.

Using it as a context manager, so the response is reliably closed — Correct: The context manager is required so that the streamed response is reliably closed.

Switching to the async client, which closes streamed responses by itself — Incorrect: The async client has the same requirement, with async methods.

Answer: C

Q6medium

Some requests in a large batch reached the expiration before they could be sent to the model. What does the team owe for them?

A.

Nothing; expired requests are not billed

B.

The full price, since they were submitted

C.

Half the price, the batch discount

D.

The price of the input tokens only

Show answer & explanation

Correct Answer: A

Nothing; expired requests are not billed — Correct: Requests that expired before reaching the model are not billed.

The full price, since they were submitted — Incorrect: Submission alone is not billed.

Half the price, the batch discount — Incorrect: The discount applies to processed requests.

The price of the input tokens only — Incorrect: No tokens were processed.

Answer: A

Q7hard

This registration makes every tool call slower, not just shell commands. The scan must be able to block a dangerous command.

python
hooks = {"PreToolUse": [ HookMatcher( hooks=[scan_shell_command]) ]}

What is the fix?

A.

Add matcher="Bash", so the scan runs for the shell calls alone

B.

Register it for PostToolUse, so it runs after each call

C.

Return an async output, so calls stop waiting on the scan

D.

Rename the callback scan_bash, so the SDK applies it to Bash

Show answer & explanation

Correct Answer: A

Add matcher="Bash", so the scan runs for the shell calls alone — Correct: A hook without a matcher runs for every event of that type; a matcher limits it to the tools it is about, and the scan can still deny.

Register it for PostToolUse, so it runs after each call — Incorrect: Moving the event does not narrow which tools trigger it, and after the call it is too late to block.

Return an async output, so calls stop waiting on the scan — Incorrect: Async outputs can’t block, modify, or inject context into the operation, since the agent has already moved on, so the scan could no longer stop a dangerous command.

Rename the callback scan_bash, so the SDK applies it to Bash — Incorrect: The SDK tests a matcher pattern against the tool name; the callback’s own name plays no part.

Answer: A

Q8medium

In an Agent SDK app, a developer registers a PreToolUse hook with the matcher ".env" to stop writes to environment files. The callback never fires. What should the developer change?

A.

Register the hook on the post-tool-use event instead, since that event sees file paths

B.

Match the file-writing tools, and check the file path inside the callback

C.

Escape the dot in the matcher, since the matcher string is read as a regular expression

D.

Return an ask decision from the callback, so the matcher is evaluated for each call

Show answer & explanation

Correct Answer: B

Register the hook on the post-tool-use event instead, since that event sees file paths — Incorrect: After the write it is too late, and matchers still test tool names.

Match the file-writing tools, and check the file path inside the callback — Correct: Matchers only match tool names, not file paths or other arguments; to filter by path, check tool_input.file_path inside the hook.

Escape the dot in the matcher, since the matcher string is read as a regular expression — Incorrect: Escaping does not help; the matcher is tested against the tool name.

Return an ask decision from the callback, so the matcher is evaluated for each call — Incorrect: A decision the callback returns cannot change whether it is called.

Answer: B

Q9easy

When the memory tool is enabled, what does Claude do before it starts a task?

A.

It deletes the memory directory to start fresh

B.

It asks the user which memory files to load

C.

It checks its memory directory for earlier context

D.

It uploads its memory files to Anthropic

Show answer & explanation

Correct Answer: C

It deletes the memory directory to start fresh — Incorrect: Memory is meant to persist, not be cleared.

It asks the user which memory files to load — Incorrect: Checking memory is automatic.

It checks its memory directory for earlier context — Correct: Claude automatically checks its memory directory before starting a task.

It uploads its memory files to Anthropic — Incorrect: Memory is client-side; nothing is uploaded to Anthropic.

Answer: C

Q10easy

After an upgrade, a prompt still says “verify every step twice”, added for an older model, and the bill rises with no gain in accuracy. What should the team do?

A.

Keep the text, so that the comparison between models stays controlled

B.

Raise effort so that the new model can follow the old text

C.

Move back to the older model the prompt was written for

D.

Audit the prompt against the current model and remove the stale text

Show answer & explanation

Correct Answer: D

Keep the text, so that the comparison between models stays controlled — Incorrect: Stale text is what drives the extra spend.

Raise effort so that the new model can follow the old text — Incorrect: Raising effort adds cost to over-specified instructions.

Move back to the older model the prompt was written for — Incorrect: The fix is the prompt, not the model.

Audit the prompt against the current model and remove the stale text — Correct: A prompt accumulates text written for a model you no longer use; auditing it against the current model is a free win.

Answer: D

Q11medium

A maths-grading service with thinking enabled wants Claude to follow the same reasoning procedure the markers use. Which technique fits?

A.

Describe the procedure once in a tool description

B.

Turn thinking off, so the procedure is followed literally

C.

Raise effort step by step until the markers’ procedure is followed on every script

D.

Put thinking tags inside the few-shot examples to show the reasoning

Show answer & explanation

Correct Answer: D

Describe the procedure once in a tool description — Incorrect: A tool description is not where a reasoning pattern is shown.

Turn thinking off, so the procedure is followed literally — Incorrect: Turning thinking off removes the reasoning you want to shape.

Raise effort step by step until the markers’ procedure is followed on every script — Incorrect: Effort changes depth, not which procedure is followed.

Put thinking tags inside the few-shot examples to show the reasoning — Correct: Use thinking tags inside few-shot examples to show Claude the reasoning pattern.

Answer: D

Q12medium

With detailed beta tracing on, a team suspects a hook is slowing down every tool call. Which span shows the hook’s own time?

A.

The span that wraps each hook execution

B.

The span that wraps each Claude API call

C.

The per-turn interaction span alone

D.

The tool span’s execution child

Show answer & explanation

Correct Answer: A

The span that wraps each hook execution — Correct: A separate span wraps each hook execution, so its time can be read on its own.

The span that wraps each Claude API call — Incorrect: API-call spans cover the model request, not hooks.

The per-turn interaction span alone — Incorrect: The turn span includes everything, so it cannot isolate the hook.

The tool span’s execution child — Incorrect: The execution child covers the tool itself.

Answer: A

Q13medium

A ticket-summary feature uses one model call. It meets the quality bar, except that about one summary in ten leaves out the account number the next team needs. Following Anthropic’s guidance, what is the smallest change that fixes this?

A.

Add a code check for the account number, and a second call that repairs summaries missing it

B.

Replace the call with an agent that decides for itself how to summarize each ticket

C.

Move the feature to orchestrator-workers, with a worker for each field of the summary it writes

D.

Keep the design as it is, and accept that one summary in ten will be incomplete

Show answer & explanation

Correct Answer: A

Add a code check for the account number, and a second call that repairs summaries missing it — Correct: A programmatic check on an intermediate step, followed by one more call, is the smallest workflow that closes the known gap, and complexity should grow only as far as needed.

Replace the call with an agent that decides for itself how to summarize each ticket — Incorrect: An agent adds latency, cost and autonomy the task does not need.

Move the feature to orchestrator-workers, with a worker for each field of the summary it writes — Incorrect: Orchestration decides subtasks per input; the fields here are fixed.

Keep the design as it is, and accept that one summary in ten will be incomplete — Incorrect: The gap is known and fixable, so accepting it leaves the bar unmet.

Answer: A

Q14easy

A developer adds a tools array to a request that already has its own system prompt. What does the API do with the tool definitions?

A.

Combines them with its tool configuration and your system prompt

B.

Replaces the developer’s system prompt with the tool definitions

C.

Sends them only to the first turn of the conversation

D.

Ignores them until the system prompt mentions each tool by name

Show answer & explanation

Correct Answer: A

Combines them with its tool configuration and your system prompt — Correct: The API constructs a special system prompt from the tool definitions, the tool configuration and any user-specified system prompt, so yours is combined, not replaced.

Replaces the developer’s system prompt with the tool definitions — Incorrect: Your system prompt is combined with the tool material, not replaced.

Sends them only to the first turn of the conversation — Incorrect: Tools apply to every request that carries them.

Ignores them until the system prompt mentions each tool by name — Incorrect: Claude sees defined tools without them being named in your prompt.

Answer: A

Q15easy

A developer wants a familiar example of a hybrid context strategy. How does Claude Code combine up-front and just-in-time context?

A.

It loads every file in the repository at the start of a session

B.

Its instruction files load up front, and glob and grep fetch files as needed

C.

It loads nothing up front and discovers even its instructions with tools

D.

It builds an embedding index of the repository before every request

Show answer & explanation

Correct Answer: B

It loads every file in the repository at the start of a session — Incorrect: Loading every file is the opposite of just in time.

Its instruction files load up front, and glob and grep fetch files as needed — Correct: Claude Code drops its instruction files into context up front, while glob and grep retrieve files just in time.

It loads nothing up front and discovers even its instructions with tools — Incorrect: Its instruction files are loaded up front, not discovered.

It builds an embedding index of the repository before every request — Incorrect: Claude Code uses glob and grep rather than a pre-built index, avoiding stale indexing.

Answer: B

These are 15 of 503 questions available. Take a practice test →

Claude Certified Developer - Foundations (CCDV-F) Flashcards

222 flashcards for spaced-repetition study. Showing 30 sample cards below.

Agent and workflow architecture — flashcards(8 cards shown)

Question

How does an agent typically begin its work?

Answer

With either a command from, or an interactive discussion with, the human user.

Question

Stripped to its mechanics, what is an agent, in Anthropic’s description?

Answer

Typically just a model using tools, based on feedback from its environment, in a loop.

Question

How can a subagent enforce constraints?

Answer

By limiting which tools the subagent can use.

Question

How can subagents help control costs?

Answer

By routing tasks to faster, cheaper models, such as Haiku.

Question

Besides coding, what kind of task does the article list for orchestrator-workers?

Answer

Search tasks that gather and analyze information from multiple sources for possibly relevant information.

Question

Does Claude Code come with subagents you do not define yourself?

Answer

Yes. It includes built-in subagents that Claude automatically uses when appropriate.

Question

What umbrella term does Anthropic use for both workflows and agents?

Answer

Agentic systems. It groups the variations together, then draws an architectural distinction between workflows and agents.

Question

How can one subagent definition serve several projects?

Answer

Define it as a user-level subagent, which reuses the configuration across projects.

Agent patterns and frameworks — flashcards(8 cards shown)

Question

Into which three buckets does every tool fall?

Answer

User-defined tools, Anthropic-schema tools, and server-executed tools. The bucket determines what your application is responsible for.

Question

When is prompt chaining ideal?

Answer

When the task can be easily and cleanly decomposed into fixed subtasks.

Question

Which kinds of task don’t need a tool round trip?

Answer

Summarization, translation and general-knowledge questions: the model can answer them from training alone.

Question

Across multiple agent sessions, what can the memory tool maintain?

Answer

Project context, so a later session continues where an earlier one stopped.

Question

Under which path does Claude store what it learns with the memory tool?

Answer

Under /memories, reading the files back in later conversations to continue earlier work.

Question

When is parallelization effective?

Answer

When the subtasks can run in parallel for speed, or when multiple perspectives or attempts are needed for higher-confidence results.

Question

What low-level tasks do agent frameworks simplify?

Answer

Calling models, defining and parsing tools, and chaining calls together.

Question

Besides hiding prompts, what other risk do frameworks bring?

Answer

They can make it tempting to add complexity when a simpler setup would suffice.

Building agents with Claude — flashcards(8 cards shown)

Question

In Managed Agents (beta), what is a session?

Answer

A running agent instance within an environment, performing a specific task and generating outputs.

Question

What kind of work is Managed Agents (beta) best for, compared with the Messages API?

Answer

Long-running tasks and asynchronous work. The Messages API suits custom agent loops and fine-grained control.

Question

In the Agent SDK, what does max_turns count?

Answer

Tool-use turns only.

Question

What happens in each cycle of the Agent SDK’s agent loop?

Answer

Claude evaluates the prompt, calls tools to take action, receives the results, and repeats until the task is complete.

Question

Which three kinds of agent state does a self-hosted SDK container keep on disk by default?

Answer

Session transcripts, CLAUDE.md memory files, and working-directory artifacts.

Question

Why is hosting an Agent SDK agent unlike hosting an API wrapper?

Answer

Each running agent is a long-lived process tied to local state, so hosting it is not like hosting a stateless API wrapper.

Question

When an Agent SDK run hits its turn cap, what does the SDK return?

Answer

A ResultMessage with the error_max_turns subtype. Hitting the budget cap returns error_max_budget_usd instead.

Question

Name three kinds of agent event a hook can respond to.

Answer

A tool being called, a session starting, and execution stopping.

Choosing models and managing cost — flashcards(6 cards shown)

Question

When you test candidate models on your own data, which three aspects does the selection guide compare?

Answer

Accuracy of responses, response quality, and handling of edge cases.

Question

Name two kinds of application the selection guide lists as best suited to a capability-first start.

Answer

For example, complex reasoning tasks, and tasks requiring nuanced understanding.

Question

Where do you look up a model’s context window, output limits and prices?

Answer

In the model comparison table. The figures change, so look them up rather than remembering them.

Question

What single question decides between the two multi-model strategies?

Answer

Does the work split into independent pieces, or is it one answer reached through a chain of dependent steps?

Question

What does a policy of re-running failures at higher effort depend on?

Answer

A failure signal. A checker that passes bad work lets those failures through.

Question

If an effort sweep leaves a quality gap, what do you price next?

Answer

The stronger model alone, at low effort. That single-model result is the baseline any multi-model strategy must beat.

Showing 30 of 222 flashcards. Study all flashcards →

Ready to ace Claude Certified Developer - Foundations (CCDV-F)?

Access all 503 practice questions, study notes, and flashcards — no sign-up required.

Start Studying — Free