Agents and Workflows — practice exercise
Difficulty: intermediate · Estimated duration: 60–90 minutes
You will design and write the core of a small agent system. The work is code and design written in your editor. No stage needs an API key or spends any credit. You compare what you write with the reference solution at the end. One optional step lets you run your loop against the live API; if you take it, it uses your own API key and your own credit.
The scenario is invented for practice. You are a developer on an accounts-receivable team. Customers dispute invoices by email, and the team wants Claude to help resolve the disputes.
| Stage | What you practise | Minutes |
|---|---|---|
| 1 | Workflow, agent or orchestrator, and subagents | 15 |
| 2 | Choosing a surface and writing the agent loop | 20 |
| 3 | Hosting the agent and guarding it with a hook | 15 |
| 4 | Tool contract, memory and workflow patterns | 20 |
Stage 1 — Choose the architecture
Skills: CCDVF-U1.T1.LO1.S1, CCDVF-U1.T1.LO1.S2, CCDVF-U1.T1.LO1.S3 Minutes: 15
The team lists three features:
- A. Each dispute email is classified as pricing, delivery or duplicate charge, and a template reply is drafted for that category. The steps never change.
- B. For a disputed invoice, find out what went wrong. Depending on the dispute this may mean checking the order system, the delivery tracker, the payment processor, or several of them. Nobody can list the checks in advance.
- C. While working on B, the agent sometimes has to read a customer’s full payment history, which runs to thousands of lines, only to find two or three relevant transactions.
- For each of A and B, write one sentence: workflow or agent (or orchestrator-workers), and why.
- B could instead be built as orchestrator-workers. Write two sentences on what the orchestrator would decide, and what each worker would do.
- For C, describe a subagent: what it reads, what tools it needs, and what it returns to the main conversation.
- Before B touches real customer accounts, scope its tools and permissions: name the tools it may only read with, and the one action it must never be able to take on its own.
Stage 2 — Pick the surface and write the loop
Skills: CCDVF-U1.T2.LO2.S1, CCDVF-U1.T2.LO2.S2 Minutes: 20
The team runs Python services on its own servers and wants full control over each call.
- Choose a surface: the Agent SDK, the Client SDK with your own loop, or Managed Agents (beta). Justify it in two sentences.
- Write the agentic loop in Python for feature B, using the Client SDK. Assume a
client, aTOOLSlist, and a functionrun_tool(block)that executes one tool call and returns its output as a string, raising an exception if the tool fails. Your loop must:- keep one running message list;
- stop when
stop_reasonis no longer"tool_use"; - answer every
tool_useblock in a response, in one user message; - return a failing tool’s error with
is_error: Trueinstead of raising; - stop after 15 turns even if Claude has not finished, and say what happens then.
- (Optional; uses your own API key and credit.) Run the loop against one stub tool that returns a fixed invoice record, and check the loop exits cleanly.
Stage 3 — Host it and guard it
Skills: CCDVF-U1.T2.LO2.S3, CCDVF-U1.T2.LO2.S4 Minutes: 15
The team is now considering moving feature B to the Agent SDK in containers.
- Disputes now arrive as a continuous stream from a busy shared inbox, all day. Choose a session pattern from the hosting guide and justify it.
- List what agent state would live on the container’s disk, and what happens to it on a scale-down.
- Sketch a Python
PreToolUsehook callback that denies any Bash command that updates theledgertable, with a reason the agent can read. Make the check robust to how people type SQL.
Stage 4 — Patterns, tools and memory
Skills: CCDVF-U1.T3.LO3.S1, CCDVF-U1.T3.LO3.S2, CCDVF-U1.T3.LO3.S3 Minutes: 20
- Claude needs to look up an invoice by its number before it can reason about a dispute. Write the tool definition: its name, a description Claude can act on, and its input schema.
- The agent should remember each customer’s past disputes between sessions, using the memory tool. Write
handle(inp, store), the part of your memory handler that dispatches each memory-tool command (view,create,str_replace,insert,delete,rename) to astoreobject. Assumestorealready maps paths onto your storage and checks them. Say what your function returns when a command fails. - Name the workflow pattern for each, in one line:
- (a) refunds above a set amount get three independent assessments, and a refund goes ahead only if at least two agree;
- (b) each proposed refund is checked at the same time by three separate calls, one for policy, one for fraud signals and one for the customer’s history, and the findings are combined;
- (c) dispute emails arrive in five languages, and each is best handled by a prompt written for that language.
Acceptance checks
| Stage | You are done when |
|---|---|
| 1 | A and B each have a design and a reason; C has a subagent with narrow tools and a summary output; B’s tools are scoped, with one action kept for a person |
| 2 | Your loop meets all five rules, handles the cap, and your surface choice names who runs the loop |
| 3 | Your pattern matches the traffic; your state list names what a scale-down loses; your hook denies with a reason, and its check survives case and spacing |
| 4 | The tool has a clear description and schema; handle covers all six commands and answers a failure with a message, not an exception; the three patterns are named |
Reference solution
Stage 1. A is a workflow: the steps are known and fixed, so predefined code paths give predictability and consistency. Classification into three categories followed by a template is also a natural fit for routing.
B needs an agent or orchestrator-workers: the checks depend on the dispute, so they cannot be hardcoded. As orchestrator-workers, the orchestrator reads the dispute and decides which systems to check. Each worker runs one check, such as querying the delivery tracker, and the orchestrator then synthesizes the findings.
For C, a subagent reads the payment history in its own context window, with read-only access to the payments tool. It returns only the relevant transactions as a short summary, so the history never floods the main conversation.
Scoping for B: read-only access to the order system, the delivery tracker and the payment processor, because investigating a dispute only needs to look. Issuing a refund or credit is not a tool it has: it recommends one, with its evidence, and a person carries it out. Scoping tools this way limits what a wrong step can do, whatever the model decides.
Stage 2. The Client SDK with your own loop fits a team that wants full control and runs its own Python services: with the Client SDK you write the tool loop yourself. The Agent SDK would also run in their process, but it brings Claude Code’s own tools and configuration, which this task does not need. A loop that meets the five rules:
USE = "tool_use"
ask = client.messages.create
def result(b):
res = {"type": "tool_result",
"tool_use_id": b.id}
try:
res["content"] = run_tool(b)
except Exception as err:
res["content"] = str(err)
res["is_error"] = True
return res
msgs = [{"role": "user",
"content": dispute}]
for _ in range(15): # turn cap
r = ask(model=MODEL,
max_tokens=1024,
tools=TOOLS,
messages=msgs)
msgs.append({
"role": "assistant",
"content": r.content})
if r.stop_reason != USE:
break
out = [result(b)
for b in r.content
if b.type == USE]
msgs.append({"role": "user",
"content": out})Each tool_result carries the tool_use_id of the call it answers. All results for one response go back in one user message, and a failing tool becomes an is_error result that Claude can react to. It can retry with corrected input, ask for clarification, or explain the limitation.
After the loop, check r.stop_reason:
end_turnis a final answer.max_tokensorrefusalneed their own handling.- If the loop ended because it hit the 15-turn cap,
stop_reasonis still"tool_use": the last tool results were appended but never sent. Treat that as an unfinished run, and log it or hand it to a person rather than presenting it as an answer.
Stage 3. Long-running sessions fit a continuous stream: persistent container instances serve ongoing work, and the guide lists high-volume message streams among their uses.
By default the container’s disk holds the session transcripts, CLAUDE.md memory files and working-directory artifacts. None of them survive a restart, scale-down or move to another node, so anything needed later must go to storage outside the container.
A hook callback, following the shape in the hooks guide:
import re
LEDGER = re.compile(
r"\bupdate\s+ledger\b", re.I)
async def block_ledger(
data, tool_use_id, ctx):
tool = data["tool_input"]
cmd = tool.get("command", "")
if not LEDGER.search(cmd):
return {}
out = dict(
hookEventName=
data["hook_event_name"],
permissionDecision="deny",
permissionDecisionReason=(
"Ledger updates need "
"finance review."))
return {"hookSpecificOutput":
out}A plain substring test such as "UPDATE ledger" in cmd misses update ledger or UPDATE ledger with two spaces. The regular expression ignores case and allows any whitespace. The reason string matters too: the agent reads it, so say what to do instead of only what is refused.
Stage 4. (1) A tool definition:
{"name": "get_invoice",
"description": (
"Look up one invoice by "
"its number. Use it before "
"judging any dispute about "
"that invoice. Returns the "
"amount, dates and status."),
"input_schema": {
"type": "object",
"properties": {
"invoice_number": {
"type": "string"}},
"required": [
"invoice_number"]}}The description says what the tool does and when to use it, and the schema names one required input, so Claude has no guessing to do.
(2) A dispatcher for the six memory commands:
def handle(inp, store):
cmd = inp["command"]
p = inp.get("path")
try:
if cmd == "view":
return store.view(p)
if cmd == "create":
store.write(
p, inp["file_text"])
return "Created " + p
if cmd == "str_replace":
return store.replace(
p, inp["old_str"],
inp["new_str"])
if cmd == "insert":
return store.insert(
p,
inp["insert_line"],
inp["insert_text"])
if cmd == "delete":
return store.delete(p)
if cmd == "rename":
return store.rename(
inp["old_path"],
inp["new_path"])
return ("Error: unknown "
"command " + cmd)
except Exception as e:
return f"Error: {e}"The memory tool is client-side: Claude requests each operation, and this function carries it out. Whatever it returns becomes the tool_result content. A failure, such as an unknown command, a missing key or a store method that raises because old_str is not in the file, comes back as a message Claude can read and correct its next call from, not as an exception that stops the loop.
(3) (a) Parallelization by voting: the same assessment runs several times, and a vote threshold decides the outcome. (b) Parallelization by sectioning: the refund check splits into independent subtasks that run at the same time, and their findings are combined. (c) Routing: each email is classified by language and directed to the prompt specialized for it.
Sources
- Building effective agents — https://www.anthropic.com/engineering/building-effective-agents
- Create custom subagents — https://code.claude.com/docs/en/sub-agents
- Agent SDK overview — https://code.claude.com/docs/en/agent-sdk/overview
- Claude Managed Agents overview — https://platform.claude.com/docs/en/managed-agents/overview
- Build a tool-using agent — https://platform.claude.com/docs/en/agents-and-tools/tool-use/build-a-tool-using-agent
- How the agent loop works — https://code.claude.com/docs/en/agent-sdk/agent-loop
- Hosting the Agent SDK — https://code.claude.com/docs/en/agent-sdk/hosting
- Intercept and control agent behavior with hooks — https://code.claude.com/docs/en/agent-sdk/hooks
- How tool use works — https://platform.claude.com/docs/en/agents-and-tools/tool-use/how-tool-use-works
- Memory tool — https://platform.claude.com/docs/en/agents-and-tools/tool-use/memory-tool





