Hands-on Lab1,954 words

Agents and Workflows — practice exercise

Agents and Workflows — practice exercise

Difficulty: intermediate · Estimated duration: 60–90 minutes

You will design and write the core of a small agent system. The work is code and design written in your editor. No stage needs an API key or spends any credit. You compare what you write with the reference solution at the end. One optional step lets you run your loop against the live API; if you take it, it uses your own API key and your own credit.

The scenario is invented for practice. You are a developer on an accounts-receivable team. Customers dispute invoices by email, and the team wants Claude to help resolve the disputes.

StageWhat you practiseMinutes
1Workflow, agent or orchestrator, and subagents15
2Choosing a surface and writing the agent loop20
3Hosting the agent and guarding it with a hook15
4Tool contract, memory and workflow patterns20

Stage 1 — Choose the architecture

Skills: CCDVF-U1.T1.LO1.S1, CCDVF-U1.T1.LO1.S2, CCDVF-U1.T1.LO1.S3 Minutes: 15

The team lists three features:

  • A. Each dispute email is classified as pricing, delivery or duplicate charge, and a template reply is drafted for that category. The steps never change.
  • B. For a disputed invoice, find out what went wrong. Depending on the dispute this may mean checking the order system, the delivery tracker, the payment processor, or several of them. Nobody can list the checks in advance.
  • C. While working on B, the agent sometimes has to read a customer’s full payment history, which runs to thousands of lines, only to find two or three relevant transactions.
  1. For each of A and B, write one sentence: workflow or agent (or orchestrator-workers), and why.
  2. B could instead be built as orchestrator-workers. Write two sentences on what the orchestrator would decide, and what each worker would do.
  3. For C, describe a subagent: what it reads, what tools it needs, and what it returns to the main conversation.
  4. Before B touches real customer accounts, scope its tools and permissions: name the tools it may only read with, and the one action it must never be able to take on its own.

Stage 2 — Pick the surface and write the loop

Skills: CCDVF-U1.T2.LO2.S1, CCDVF-U1.T2.LO2.S2 Minutes: 20

The team runs Python services on its own servers and wants full control over each call.

  1. Choose a surface: the Agent SDK, the Client SDK with your own loop, or Managed Agents (beta). Justify it in two sentences.
  2. Write the agentic loop in Python for feature B, using the Client SDK. Assume a client, a TOOLS list, and a function run_tool(block) that executes one tool call and returns its output as a string, raising an exception if the tool fails. Your loop must:
    • keep one running message list;
    • stop when stop_reason is no longer "tool_use";
    • answer every tool_use block in a response, in one user message;
    • return a failing tool’s error with is_error: True instead of raising;
    • stop after 15 turns even if Claude has not finished, and say what happens then.
  3. (Optional; uses your own API key and credit.) Run the loop against one stub tool that returns a fixed invoice record, and check the loop exits cleanly.

Stage 3 — Host it and guard it

Skills: CCDVF-U1.T2.LO2.S3, CCDVF-U1.T2.LO2.S4 Minutes: 15

The team is now considering moving feature B to the Agent SDK in containers.

  1. Disputes now arrive as a continuous stream from a busy shared inbox, all day. Choose a session pattern from the hosting guide and justify it.
  2. List what agent state would live on the container’s disk, and what happens to it on a scale-down.
  3. Sketch a Python PreToolUse hook callback that denies any Bash command that updates the ledger table, with a reason the agent can read. Make the check robust to how people type SQL.

Stage 4 — Patterns, tools and memory

Skills: CCDVF-U1.T3.LO3.S1, CCDVF-U1.T3.LO3.S2, CCDVF-U1.T3.LO3.S3 Minutes: 20

  1. Claude needs to look up an invoice by its number before it can reason about a dispute. Write the tool definition: its name, a description Claude can act on, and its input schema.
  2. The agent should remember each customer’s past disputes between sessions, using the memory tool. Write handle(inp, store), the part of your memory handler that dispatches each memory-tool command (view, create, str_replace, insert, delete, rename) to a store object. Assume store already maps paths onto your storage and checks them. Say what your function returns when a command fails.
  3. Name the workflow pattern for each, in one line:
    • (a) refunds above a set amount get three independent assessments, and a refund goes ahead only if at least two agree;
    • (b) each proposed refund is checked at the same time by three separate calls, one for policy, one for fraud signals and one for the customer’s history, and the findings are combined;
    • (c) dispute emails arrive in five languages, and each is best handled by a prompt written for that language.

Acceptance checks

StageYou are done when
1A and B each have a design and a reason; C has a subagent with narrow tools and a summary output; B’s tools are scoped, with one action kept for a person
2Your loop meets all five rules, handles the cap, and your surface choice names who runs the loop
3Your pattern matches the traffic; your state list names what a scale-down loses; your hook denies with a reason, and its check survives case and spacing
4The tool has a clear description and schema; handle covers all six commands and answers a failure with a message, not an exception; the three patterns are named

Reference solution

Stage 1. A is a workflow: the steps are known and fixed, so predefined code paths give predictability and consistency. Classification into three categories followed by a template is also a natural fit for routing.

B needs an agent or orchestrator-workers: the checks depend on the dispute, so they cannot be hardcoded. As orchestrator-workers, the orchestrator reads the dispute and decides which systems to check. Each worker runs one check, such as querying the delivery tracker, and the orchestrator then synthesizes the findings.

For C, a subagent reads the payment history in its own context window, with read-only access to the payments tool. It returns only the relevant transactions as a short summary, so the history never floods the main conversation.

Scoping for B: read-only access to the order system, the delivery tracker and the payment processor, because investigating a dispute only needs to look. Issuing a refund or credit is not a tool it has: it recommends one, with its evidence, and a person carries it out. Scoping tools this way limits what a wrong step can do, whatever the model decides.

Stage 2. The Client SDK with your own loop fits a team that wants full control and runs its own Python services: with the Client SDK you write the tool loop yourself. The Agent SDK would also run in their process, but it brings Claude Code’s own tools and configuration, which this task does not need. A loop that meets the five rules:

python
USE = "tool_use" ask = client.messages.create def result(b): res = {"type": "tool_result", "tool_use_id": b.id} try: res["content"] = run_tool(b) except Exception as err: res["content"] = str(err) res["is_error"] = True return res msgs = [{"role": "user", "content": dispute}] for _ in range(15): # turn cap r = ask(model=MODEL, max_tokens=1024, tools=TOOLS, messages=msgs) msgs.append({ "role": "assistant", "content": r.content}) if r.stop_reason != USE: break out = [result(b) for b in r.content if b.type == USE] msgs.append({"role": "user", "content": out})

Each tool_result carries the tool_use_id of the call it answers. All results for one response go back in one user message, and a failing tool becomes an is_error result that Claude can react to. It can retry with corrected input, ask for clarification, or explain the limitation.

After the loop, check r.stop_reason:

  • end_turn is a final answer.
  • max_tokens or refusal need their own handling.
  • If the loop ended because it hit the 15-turn cap, stop_reason is still "tool_use": the last tool results were appended but never sent. Treat that as an unfinished run, and log it or hand it to a person rather than presenting it as an answer.

Stage 3. Long-running sessions fit a continuous stream: persistent container instances serve ongoing work, and the guide lists high-volume message streams among their uses.

By default the container’s disk holds the session transcripts, CLAUDE.md memory files and working-directory artifacts. None of them survive a restart, scale-down or move to another node, so anything needed later must go to storage outside the container.

A hook callback, following the shape in the hooks guide:

python
import re LEDGER = re.compile( r"\bupdate\s+ledger\b", re.I) async def block_ledger( data, tool_use_id, ctx): tool = data["tool_input"] cmd = tool.get("command", "") if not LEDGER.search(cmd): return {} out = dict( hookEventName= data["hook_event_name"], permissionDecision="deny", permissionDecisionReason=( "Ledger updates need " "finance review.")) return {"hookSpecificOutput": out}

A plain substring test such as "UPDATE ledger" in cmd misses update ledger or UPDATE ledger with two spaces. The regular expression ignores case and allows any whitespace. The reason string matters too: the agent reads it, so say what to do instead of only what is refused.

Stage 4. (1) A tool definition:

python
{"name": "get_invoice", "description": ( "Look up one invoice by " "its number. Use it before " "judging any dispute about " "that invoice. Returns the " "amount, dates and status."), "input_schema": { "type": "object", "properties": { "invoice_number": { "type": "string"}}, "required": [ "invoice_number"]}}

The description says what the tool does and when to use it, and the schema names one required input, so Claude has no guessing to do.

(2) A dispatcher for the six memory commands:

python
def handle(inp, store): cmd = inp["command"] p = inp.get("path") try: if cmd == "view": return store.view(p) if cmd == "create": store.write( p, inp["file_text"]) return "Created " + p if cmd == "str_replace": return store.replace( p, inp["old_str"], inp["new_str"]) if cmd == "insert": return store.insert( p, inp["insert_line"], inp["insert_text"]) if cmd == "delete": return store.delete(p) if cmd == "rename": return store.rename( inp["old_path"], inp["new_path"]) return ("Error: unknown " "command " + cmd) except Exception as e: return f"Error: {e}"

The memory tool is client-side: Claude requests each operation, and this function carries it out. Whatever it returns becomes the tool_result content. A failure, such as an unknown command, a missing key or a store method that raises because old_str is not in the file, comes back as a message Claude can read and correct its next call from, not as an exception that stops the loop.

(3) (a) Parallelization by voting: the same assessment runs several times, and a vote threshold decides the outcome. (b) Parallelization by sectioning: the refund check splits into independent subtasks that run at the same time, and their findings are combined. (c) Routing: each email is classified by language and directed to the prompt specialized for it.

Sources

Ready to study Claude Certified Developer - Foundations (CCDV-F)?

Practice tests, flashcards, and all study notes — free, no sign-up needed.

Start Studying — Free