Prompt and Context Engineering — study note
Prompt and Context Engineering — study note
This note covers Domain 6 of the Claude Certified Developer – Foundations exam. It has three skills:
- context engineering;
- prompt engineering;
- output handling.
It summarizes what each topic teaches and the Anthropic pages behind it. It teaches mechanisms rather than numbers: model names, limits and beta identifiers change often, so look them up on the linked pages when you need them.
Context is a budget you curate
Context engineering is the natural progression of prompt engineering. It curates the whole set of tokens the model sees, not only the prompt, and the curation happens again every time you decide what to pass to the model.
- Finite budget. Every new token depletes the model's attention budget. Longer contexts show a gradient, not a cliff: the model stays capable, but precision for retrieval and long-range reasoning can drop.
- Guiding principle. Find the smallest set of high-signal tokens that maximize the likelihood of the outcome you want.
| Loading strategy | Trade-off |
|---|---|
| Up front | Fast, but loads material a task may not need |
| Just in time | Keeps identifiers and loads data with tools; exploration is slower |
| Hybrid | Some data up front for speed, the rest explored as needed |
Claude Code is a hybrid: its instruction files load up front, while glob and grep retrieve files just in time.
Compaction and context editing
Compaction replaces older turns with a summary that Claude writes on the server.
| Compaction | Who decides when | Choose it when |
|---|---|---|
| On demand | You, by sending a request | You control timing, can't pause, or keep recent turns |
| At a threshold | The API, at your trigger | You want the API to manage it inside ordinary requests |
| Your own summarizer | You | You already run one that replaces the whole history |
Use on-demand compaction wherever it is available. Its options can keep recent turns word for word or run in the background, and you can write your own summarization prompt. Tune a custom compaction prompt on complex agent traces: maximise recall first, then trim for precision.
Context editing clears content by rule instead of summarizing it. Tool result clearing removes the oldest results first, keeps the most recent ones, and leaves a placeholder for each.
| Setting | Effect |
|---|---|
keep | How many recent tool uses survive |
exclude_tools | Tools whose results are never cleared |
clear_at_least | Skip a pass that cannot clear this much |
- Clearing tool results invalidates the cached prefix, so clear enough to make it worthwhile.
- Setting
clear_tool_inputsto true clears tool call parameters as well as results. - Editing happens server-side: your client keeps the full history and needs no syncing. The
context_managementresponse field reports what was cleared, and the token counting endpoint can preview the count after editing. - With the memory tool, Claude is warned as clearing approaches, so it can save findings first.
- Tell Claude about your harness. If context is compacted or saved to files, say so in the prompt, or Claude may wrap up early near the limit.
- Across windows. Use the first window to set up tests and setup scripts; later windows can start fresh and read the progress file, the tests file and the git log.
Compaction suits long research and multistep work with measurable progress. It is less suited to tasks that need precise recall of early details or exact state across many variables. Anthropic recommends server-side compaction over SDK compaction; SDK compaction can miscount usage when server-side tools are involved.
Keeping tool output and side work out of the thread
| Where tokens go | Approach |
|---|---|
| Unused tool definitions | Tool search |
| Round trips in a chain of calls | Programmatic tool calling |
| Old results in history | Context editing |
| Repeated definition cost | Prompt caching (cost, not size) |
Measure first, then pick the approach aimed at that source; the approaches compose without conflict. With programmatic tool calling, Claude writes one script that Anthropic's code execution sandbox runs, and the intermediate results never enter the conversation. For large data, have the agent write targeted queries, store results and inspect them with commands like head and tail.
| Keep in the main conversation | Hand to a subagent |
|---|---|
| Phases share significant context | Verbose output you won't need |
| Latency matters | Self-contained work that returns a summary |
| Frequent back-and-forth | Specific tool restrictions |
A subagent starts with a fresh context: it doesn't see the conversation history, and works from the delegation message Claude writes. Its window is sized by its own model. For a quick question about something already in the conversation, /btw sees the full context without adding to history. For long-horizon work, match the technique to the task: compaction for back-and-forth, structured notes for iterative work with milestones, multi-agent for parallel research.
Prompt engineering
- Is a prompt the fix? Have success criteria, a way to test against them, and a first draft. If there is no draft, generate one with the metaprompt recipe. Not every failing eval is a prompt problem: latency and cost can sometimes be fixed more easily by choosing a different model.
- Start minimal. Test a minimal prompt with the best model available, then add instructions and examples for the failures you find. Minimal is not the same as short.
- Model-specific guidance. Read your model's own prompting page first. Treat a technique that names a model as measured on that model, and re-check it on your evals.
| Instruction problem | Fix |
|---|---|
| "Can you suggest changes?" gets suggestions only | Say "Change this function…" |
| Agent should only advise | Default to recommendations; act on explicit request |
| Old CAPS emphasis now overtriggers | Dial back to "Use this tool when…" |
| 60 brittle if-then rules | Right altitude: guidance plus heuristics |
| Tests passed by special-casing | Ask for general solutions; report bad tests |
If operators need to see an agent's work, ask for a quick summary after each tool-using task. If an interrupted response has no user-facing cost, retry the request. For a pipeline you want to inspect, chain calls: generate, review against criteria, refine, with each step logged. Judge every change against a baseline on the same fixed test set.
Output handling
Structured outputs guarantee schema compliance in most cases, not all.
| Situation | Defensive move |
|---|---|
refusal stop reason | Output may not match; do not parse blindly |
max_tokens stop | Output may be incomplete; retry with more |
| Enum capitalisation differs | Compare case-insensitively |
| Optional fields come last | Read by name, or mark all required |
| SDK removed a constraint | Keep validating the response |
- Schema too complex to compile. Mark only critical tools as strict, make parameters required where possible, and flatten nesting.
- Scope. The grammar applies only to Claude's direct output, not to thinking, so Claude can think freely. Changing
output_config.formatinvalidates the prompt cache for that thread. Citations can't be combined with a JSON output format, which returns a 400. - Without a schema. For a fixed label set, use a tool with an enum field or structured outputs. For a plain-text channel, ask for plain text only and math in standard characters, since current models default to LaTeX. For formats JSON can't express, ask the model to conform and retry on failure.
- Skepticism. Tell coding agents never to speculate about code they have not opened. Mark removed, unsupported claims with empty [] brackets so code can find them. Validate critical information.
Sources
- Effective context engineering for AI agents — https://www.anthropic.com/engineering/effective-context-engineering-for-ai-agents
- Compaction overview — https://platform.claude.com/docs/en/build-with-claude/compaction
- Context editing — https://platform.claude.com/docs/en/build-with-claude/context-editing
- Manage tool context — https://platform.claude.com/docs/en/agents-and-tools/tool-use/manage-tool-context
- Create custom subagents — https://code.claude.com/docs/en/sub-agents
- Building effective agents — https://www.anthropic.com/engineering/building-effective-agents
- Prompt engineering overview — https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/overview
- Prompting best practices — https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/claude-prompting-best-practices
- Define success criteria and build evaluations — https://platform.claude.com/docs/en/test-and-evaluate/develop-tests
- Structured outputs — https://platform.claude.com/docs/en/build-with-claude/structured-outputs
- Increase output consistency — https://platform.claude.com/docs/en/test-and-evaluate/strengthen-guardrails/increase-consistency
- Reduce hallucinations — https://platform.claude.com/docs/en/test-and-evaluate/strengthen-guardrails/reduce-hallucinations