Prompt and Context Engineering — practice exercise
Prompt and Context Engineering — practice exercise
Difficulty: intermediate · Estimated duration: 45–60 minutes
You are the engineer responsible for ShelfSense, an invented inventory-planning agent. It reads stock files, calls supplier tools, writes reorder plans and runs for hours at a time. You will review its context handling, its prompts and its output parsing, and write down the decision you would make at each step. Every stage can be completed without an API key and without spending anything: you read the material given here, write code or a decision, and check it against the reference solution. Stage 3 has an optional step: if you have your own API key you may run your chain, which uses your own credit.
What you need: a text editor and a notes file. Python is used for the code; any language is fine for the reasoning.
| Stage | What you practise | Minutes |
|---|---|---|
| 1 | Reading a context report, curating context | 12 |
| 2 | Containing tool output and side work | 12 |
| 3 | Improving the prompt against a fixed set | 14 |
| 4 | Handling structured and unstructured output | 12 |
Stage 1 — Read the context report
Skills: CCDVF-U6.T1.LO1.S1, CCDVF-U6.T1.LO1.S2 Minutes: 12
ShelfSense uses tool result clearing. One response carried this report (the numbers are invented for the exercise):
{"applied_edits": [{
"type": "<clearing strategy>",
"cleared_tool_uses": 12,
"cleared_input_tokens": 48000
}]}- Which response field carries this report, and where does a streaming client find the same report?
- Before sending its next request, the team wants to know how many tokens the prompt will use once editing is applied. How can it find out without sending the request?
- A colleague argues: "We are nowhere near the window limit, so answer quality can't be affected by length." How do you reply?
- ShelfSense uses SDK compaction in the client, and a supplier-search step uses a server-side tool. Compaction fires far too early. What do you recommend?
Stage 2 — Contain tool output and side work
Skills: CCDVF-U6.T1.LO1.S3, CCDVF-U6.T1.LO1.S4 Minutes: 12
- ShelfSense's old
write_plancalls carry whole plan files in their inputs, and clearing results alone leaves those inputs in context. Write the clearing settings as a Python dict. Keep the strategy name as a constant,CLEAR_TOOLS. - Only the audit subagent uses the supplier-catalogue MCP server, and its tool descriptions crowd the main conversation. Where should the server be defined?
- An engineer wants a quick answer about a decision already discussed in the session, without adding the exchange to history. What should they use?
- The audit subagent runs on a model with a smaller window than the main session. How much context does it get?
Stage 3 — Improve the prompt against a fixed set
Skills: CCDVF-U6.T2.LO2.S1, CCDVF-U6.T2.LO2.S2, CCDVF-U6.T2.LO2.S3 Minutes: 14
- ShelfSense still explores far more supplier data than a plan needs, even after its blanket instructions were replaced with targeted ones. What is the fallback lever?
- It leaves helper scripts it wrote for iteration scattered through the working directory. Write the instruction you would add.
- Plans sometimes omit safety stock. Write a self-correction chain in Python as three separate calls (draft, review against criteria, refine) that logs each step. Assume a helper
call(prompt)that returns text, andlog(step, text). - ShelfSense's long runs span several context windows, and it has started deleting tests it cannot pass. What two things should its prompt set up?
- (Optional, uses your own API key and credit.) Implement
callwith the SDK and run your chain once on one invented stock file.
Stage 4 — Handle the output
Skills: CCDVF-U6.T2.LO3.S1, CCDVF-U6.T2.LO3.S2, CCDVF-U6.T2.LO3.S3 Minutes: 12
- ShelfSense adds strict mode to fourteen tools and a large JSON output schema, and requests now fail with "Schema is too complex for compilation". List three fixes, in order.
- The team plans to switch the output schema from turn to turn inside one cached conversation. What does that do to the prompt cache?
- ShelfSense tracks each test's status in prose and later misreads its own counts. Which format suits that state, and what suits general progress notes?
- A supplier-risk subtask writes confident figures that each rest on a single source. Write two lines to add to its prompt.
Acceptance checks
- Stage 1 names the report field and the streaming event, uses the token counting endpoint to preview, describes a gradient rather than a cliff, and moves to server-side compaction or token counting for the early compaction.
- Stage 2 clears tool inputs, defines the MCP server inline in the subagent, uses
/btw, and sizes the subagent's window by its own model. - Stage 3 names effort as the fallback, writes a clean-up instruction, logs all three calls, and protects the tests.
- Stage 4 answers with mechanisms, never with a price or a model name.
Reference solution
Stage 1. (1) The context_management field of the response, with statistics about what was cleared; for a streaming response, the report arrives in the final message_delta event. (2) Call the token counting endpoint with the same context management settings: it previews how many tokens the prompt will use after editing. (3) Length degrades quality as a gradient, not a cliff. The model stays highly capable at longer contexts, but precision for retrieval and long-range reasoning can drop, so headroom is not proof of quality. (4) With server-side tools, the SDK can miscount usage, because cache reads from the tool's internal calls inflate its total, and it compacts at the wrong time. Use the token counting endpoint for the real context length, or avoid SDK compaction there. Anthropic recommends server-side compaction over SDK compaction.
Stage 2. (1) A good answer:
edits = [{
"type": CLEAR_TOOLS,
"clear_tool_inputs": True,
}]By default only tool results are cleared; setting clear_tool_inputs to true clears the tool call parameters as well. (2) Inline in the audit subagent's definition. That keeps the server out of the main conversation, so its tool descriptions don't consume context there. (3) /btw. It sees the full context, has no tool access, and its answer isn't added to history. (4) The window of its own model: a subagent's context window is sized by its own model, not the parent's.
Stage 3. (1) Lower the effort setting. When Claude remains overly aggressive after the prompt changes, effort is the documented fallback. (2) For example: "If you create any temporary files, scripts or helper files for iteration, remove them at the end of the task." (3) A good chain:
def plan(stock):
draft = call(DRAFT + stock)
log("draft", draft)
review = call(REVIEW + draft)
log("review", review)
final = call(
REFINE + draft + review)
log("refine", final)
return finalEach step is a separate API call, so each can be logged, evaluated or branched on. When a plan still fails, the logs show which step went wrong. (4) Ask ShelfSense to create its tests before starting and keep them in a structured file such as tests.json, and state that removing or editing tests is unacceptable because it could lead to missing or buggy functionality. (5) The chain returns the refined plan and three log lines.
Stage 4. (1) Mark only the critical tools as strict, and rely on Claude's natural adherence for the rest. Reduce optional parameters, making them required where a sensible default exists. Flatten deeply nested structures. (2) Changing output_config.format invalidates the prompt cache for that conversation thread, so a per-turn schema switch loses the cache. (3) JSON or another structured format for test status. Freeform text suits general progress notes. (4) For example: "Verify each figure across multiple sources." and "Develop competing hypotheses and track your confidence in your progress notes."
Sources
- Effective context engineering for AI agents — https://www.anthropic.com/engineering/effective-context-engineering-for-ai-agents
- Context editing — https://platform.claude.com/docs/en/build-with-claude/context-editing
- Create custom subagents — https://code.claude.com/docs/en/sub-agents
- Prompting best practices — https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/claude-prompting-best-practices
- Structured outputs — https://platform.claude.com/docs/en/build-with-claude/structured-outputs