Debugging and Error Handling — practice exercise
Debugging and Error Handling — practice exercise
Difficulty: intermediate · Estimated duration: 45–60 minutes
You will diagnose a set of failures from an invented application and write the code or configuration that handles each one. The failures are given as logs and response fragments, so no API spend is needed. You can write the code in Python or TypeScript. Running any of it against the live API is optional and uses your own key.
The application: an invoice-processing service that calls the Claude API, uses one server tool (web search) and two client tools (lookup_customer, post_ledger_entry), and runs a nightly agent job through the Agent SDK with tracing enabled.
| Stage | What you practise | Minutes |
|---|---|---|
| 1 | Classifying HTTP errors and choosing retry, fix or escalate | 12 |
| 2 | Handling the stop reasons that need code | 14 |
| 3 | Placing tool failures in the integration or the model | 12 |
| 4 | Reading a trace | 12 |
Stage 1 — Error triage
Skills: CCDVF-U4.T1.LO1.S1 Minutes: 12
For each response, write the class of error and what the client should do.
401withauthentication_error, on every call since a deploy.529withoverloaded_error, on a few calls during the afternoon peak.402withbilling_error, on every call since this morning.409withconflict_error, on an update to a resource another job changed first.
Then write the error-handling block for one call. It must catch the SDK's typed exception for status errors, branch on the status code, log the request ID, and leave transient retries to the SDK.
Stage 2 — Stop reasons
Skills: CCDVF-U4.T1.LO1.S2 Minutes: 14
- Write a function that takes a successful response and handles
end_turn,max_tokens,pause_turnandrefusal. - Explain in one sentence why the function is called on responses whose HTTP status is 200.
- A response stops with
tool_useforlookup_customer. Say how this differs frompause_turn, and what your code sends next.
Stage 3 — Integration or model?
Skills: CCDVF-U4.T1.LO1.S3 Minutes: 12
For each symptom, say whether the fault is in the integration or in what the model was given, and write the fix.
- A 400 says
tool_useids were found withouttool_resultblocks immediately after. Your code sent the two results in two separate messages, each after a line of text. - Claude calls
post_ledger_entrywhen you expectedlookup_customer. - Claude sometimes passes
status: "archived"tolookup_customer, a value outside the tool'sstatusenum of forty values. - Your code puts the instruction "only look up active customers" inside the
lookup_customerresult, and Claude asks the user to confirm it.
Stage 4 — Read a trace
Skills: CCDVF-U4.T1.LO1.S4 Minutes: 12
The nightly job's trace for one turn shows:
- the turn span: 94 seconds;
- two Claude API call spans: 3 and 4 seconds;
- one
post_ledger_entrytool span: 86 seconds, whose permission-wait child is 85 seconds and whose execution child is 1 second.
- Where did the time go, and what would you change?
- The team added a hook that runs before each tool call. With detailed tracing on, which span shows the hook's own time?
- The job's spans appear as a separate trace from the application that started them. What connects them?
Acceptance checks
- Every Stage 1 answer names the class and the action, and the code catches the SDK's status-error type, branches on the status code, and logs the request ID.
- The stop-reason function branches on
stop_reason, marks truncations as incomplete, resends onpause_turn, and routes refusals tostop_detailsand a fallback model. - Each Stage 3 fix is placed correctly: integration, description, schema, or where the instructions go.
- The trace answers use what each span covers, not guesses.
Reference solution
Stage 1
| Response | Class and action |
|---|---|
| 401 | The key: fix the credential, and do not retry |
| 529 | Transient overload: retry with backoff |
| 402 | Billing: fix the payment details, and do not retry |
| 409 | Conflict: resolve it, then retry |
from anthropic import APIStatusError
try:
msg = client.messages.create(
**request)
except APIStatusError as e:
rid = e.response.headers.get(
"request-id")
log.error("status %s id %s",
e.status_code, rid)
code = e.status_code
if code in (401, 402, 403):
alert_owner() # fix access
raise
if code == 409:
msg = reload_and_retry()
else:
raiseThe SDK already retries connection errors, rate limits and 5xx errors with backoff, so the handler branches on status_code and re-raises rather than looping. 401, 402 and 403 alert the owner to fix the key, billing or access; 409 resolves the conflict before its retry. anthropic.APIStatusError and status_code are the names the stop-reasons page uses in Python, and the Python SDK page reads the ID with response.headers.get("request-id"). Error bodies carry the same ID as request_id. The helper names are placeholders.
Stage 2
NOTE = "\n[output incomplete]"
def handle(msg, history):
reason = msg.stop_reason
if reason == "max_tokens":
return text_of(msg) + NOTE
if reason == "pause_turn":
history.append({
"role": "assistant",
"content": msg.content})
return None # loop resends
if reason == "refusal":
log.info("refusal: %s",
msg.stop_details)
return None # try fallback
return text_of(msg) # end_turnA refusal arrives as a normal HTTP 200, so the status alone would pass it on as an answer; the documented action is to read stop_details and retry on a fallback model. A tool_use stop waits on your own tool, so your code runs lookup_customer and sends a tool_result. pause_turn is a server-tool loop pausing, so you resend the response as-is.
Stage 3
- Integration: send one message whose content is both
tool_resultblocks, before any text. - The model's input: sharpen the tool descriptions so they say when to use each tool.
- The model's input: shrink the enum, or add
input_examplesshowing the valid choices. - Integration: move the instruction out of the tool result and keep the result to data. The page names two places: a
userturn after thetool_resultblock, or, on supported models, a mid-conversation system message.
Stage 4
- 85 of the 86 seconds were the permission wait, not the tool. Pre-approve the call for the nightly job, or run it in a mode suited to unattended work.
- The span that wraps each hook execution; it needs detailed beta tracing.
- Start the agent run while an application span is active. The SDK propagates W3C trace context, so the run appears inside the application's trace.
Sources
Claude API errors; handling stop reasons; troubleshooting tool use; the Python SDK; Agent SDK observability.