Study Guide1,400 words

Prompt and Context Engineering — study note

Prompt and Context Engineering — study note

This note covers Domain 6 of the Claude Certified Developer – Foundations exam. It has three skills:

  • context engineering;
  • prompt engineering;
  • output handling.

It summarizes what each topic teaches and the Anthropic pages behind it. It teaches mechanisms rather than numbers: model names, limits and beta identifiers change often, so look them up on the linked pages when you need them.

Context is a budget you curate

Context engineering is the natural progression of prompt engineering. It curates the whole set of tokens the model sees, not only the prompt, and the curation happens again every time you decide what to pass to the model.

  • Finite budget. Every new token depletes the model's attention budget. Longer contexts show a gradient, not a cliff: the model stays capable, but precision for retrieval and long-range reasoning can drop.
  • Guiding principle. Find the smallest set of high-signal tokens that maximize the likelihood of the outcome you want.
Loading strategyTrade-off
Up frontFast, but loads material a task may not need
Just in timeKeeps identifiers and loads data with tools; exploration is slower
HybridSome data up front for speed, the rest explored as needed

Claude Code is a hybrid: its instruction files load up front, while glob and grep retrieve files just in time.

Compaction and context editing

Compaction replaces older turns with a summary that Claude writes on the server.

CompactionWho decides whenChoose it when
On demandYou, by sending a requestYou control timing, can't pause, or keep recent turns
At a thresholdThe API, at your triggerYou want the API to manage it inside ordinary requests
Your own summarizerYouYou already run one that replaces the whole history

Use on-demand compaction wherever it is available. Its options can keep recent turns word for word or run in the background, and you can write your own summarization prompt. Tune a custom compaction prompt on complex agent traces: maximise recall first, then trim for precision.

Context editing clears content by rule instead of summarizing it. Tool result clearing removes the oldest results first, keeps the most recent ones, and leaves a placeholder for each.

SettingEffect
keepHow many recent tool uses survive
exclude_toolsTools whose results are never cleared
clear_at_leastSkip a pass that cannot clear this much
  • Clearing tool results invalidates the cached prefix, so clear enough to make it worthwhile.
  • Setting clear_tool_inputs to true clears tool call parameters as well as results.
  • Editing happens server-side: your client keeps the full history and needs no syncing. The context_management response field reports what was cleared, and the token counting endpoint can preview the count after editing.
  • With the memory tool, Claude is warned as clearing approaches, so it can save findings first.
  • Tell Claude about your harness. If context is compacted or saved to files, say so in the prompt, or Claude may wrap up early near the limit.
  • Across windows. Use the first window to set up tests and setup scripts; later windows can start fresh and read the progress file, the tests file and the git log.

Compaction suits long research and multistep work with measurable progress. It is less suited to tasks that need precise recall of early details or exact state across many variables. Anthropic recommends server-side compaction over SDK compaction; SDK compaction can miscount usage when server-side tools are involved.

Keeping tool output and side work out of the thread

Where tokens goApproach
Unused tool definitionsTool search
Round trips in a chain of callsProgrammatic tool calling
Old results in historyContext editing
Repeated definition costPrompt caching (cost, not size)

Measure first, then pick the approach aimed at that source; the approaches compose without conflict. With programmatic tool calling, Claude writes one script that Anthropic's code execution sandbox runs, and the intermediate results never enter the conversation. For large data, have the agent write targeted queries, store results and inspect them with commands like head and tail.

Keep in the main conversationHand to a subagent
Phases share significant contextVerbose output you won't need
Latency mattersSelf-contained work that returns a summary
Frequent back-and-forthSpecific tool restrictions

A subagent starts with a fresh context: it doesn't see the conversation history, and works from the delegation message Claude writes. Its window is sized by its own model. For a quick question about something already in the conversation, /btw sees the full context without adding to history. For long-horizon work, match the technique to the task: compaction for back-and-forth, structured notes for iterative work with milestones, multi-agent for parallel research.

Prompt engineering

  • Is a prompt the fix? Have success criteria, a way to test against them, and a first draft. If there is no draft, generate one with the metaprompt recipe. Not every failing eval is a prompt problem: latency and cost can sometimes be fixed more easily by choosing a different model.
  • Start minimal. Test a minimal prompt with the best model available, then add instructions and examples for the failures you find. Minimal is not the same as short.
  • Model-specific guidance. Read your model's own prompting page first. Treat a technique that names a model as measured on that model, and re-check it on your evals.
Instruction problemFix
"Can you suggest changes?" gets suggestions onlySay "Change this function…"
Agent should only adviseDefault to recommendations; act on explicit request
Old CAPS emphasis now overtriggersDial back to "Use this tool when…"
60 brittle if-then rulesRight altitude: guidance plus heuristics
Tests passed by special-casingAsk for general solutions; report bad tests

If operators need to see an agent's work, ask for a quick summary after each tool-using task. If an interrupted response has no user-facing cost, retry the request. For a pipeline you want to inspect, chain calls: generate, review against criteria, refine, with each step logged. Judge every change against a baseline on the same fixed test set.

Output handling

Structured outputs guarantee schema compliance in most cases, not all.

SituationDefensive move
refusal stop reasonOutput may not match; do not parse blindly
max_tokens stopOutput may be incomplete; retry with more
Enum capitalisation differsCompare case-insensitively
Optional fields come lastRead by name, or mark all required
SDK removed a constraintKeep validating the response
  • Schema too complex to compile. Mark only critical tools as strict, make parameters required where possible, and flatten nesting.
  • Scope. The grammar applies only to Claude's direct output, not to thinking, so Claude can think freely. Changing output_config.format invalidates the prompt cache for that thread. Citations can't be combined with a JSON output format, which returns a 400.
  • Without a schema. For a fixed label set, use a tool with an enum field or structured outputs. For a plain-text channel, ask for plain text only and math in standard characters, since current models default to LaTeX. For formats JSON can't express, ask the model to conform and retry on failure.
  • Skepticism. Tell coding agents never to speculate about code they have not opened. Mark removed, unsupported claims with empty [] brackets so code can find them. Validate critical information.

Sources

Ready to study Claude Certified Developer - Foundations (CCDV-F)?

Practice tests, flashcards, and all study notes — free, no sign-up needed.

Start Studying — Free