Hands-on Lab1,106 words

Debugging and Error Handling — practice exercise

Debugging and Error Handling — practice exercise

Difficulty: intermediate · Estimated duration: 45–60 minutes

You will diagnose a set of failures from an invented application and write the code or configuration that handles each one. The failures are given as logs and response fragments, so no API spend is needed. You can write the code in Python or TypeScript. Running any of it against the live API is optional and uses your own key.

The application: an invoice-processing service that calls the Claude API, uses one server tool (web search) and two client tools (lookup_customer, post_ledger_entry), and runs a nightly agent job through the Agent SDK with tracing enabled.

StageWhat you practiseMinutes
1Classifying HTTP errors and choosing retry, fix or escalate12
2Handling the stop reasons that need code14
3Placing tool failures in the integration or the model12
4Reading a trace12

Stage 1 — Error triage

Skills: CCDVF-U4.T1.LO1.S1 Minutes: 12

For each response, write the class of error and what the client should do.

  1. 401 with authentication_error, on every call since a deploy.
  2. 529 with overloaded_error, on a few calls during the afternoon peak.
  3. 402 with billing_error, on every call since this morning.
  4. 409 with conflict_error, on an update to a resource another job changed first.

Then write the error-handling block for one call. It must catch the SDK's typed exception for status errors, branch on the status code, log the request ID, and leave transient retries to the SDK.

Stage 2 — Stop reasons

Skills: CCDVF-U4.T1.LO1.S2 Minutes: 14

  1. Write a function that takes a successful response and handles end_turn, max_tokens, pause_turn and refusal.
  2. Explain in one sentence why the function is called on responses whose HTTP status is 200.
  3. A response stops with tool_use for lookup_customer. Say how this differs from pause_turn, and what your code sends next.

Stage 3 — Integration or model?

Skills: CCDVF-U4.T1.LO1.S3 Minutes: 12

For each symptom, say whether the fault is in the integration or in what the model was given, and write the fix.

  1. A 400 says tool_use ids were found without tool_result blocks immediately after. Your code sent the two results in two separate messages, each after a line of text.
  2. Claude calls post_ledger_entry when you expected lookup_customer.
  3. Claude sometimes passes status: "archived" to lookup_customer, a value outside the tool's status enum of forty values.
  4. Your code puts the instruction "only look up active customers" inside the lookup_customer result, and Claude asks the user to confirm it.

Stage 4 — Read a trace

Skills: CCDVF-U4.T1.LO1.S4 Minutes: 12

The nightly job's trace for one turn shows:

  • the turn span: 94 seconds;
  • two Claude API call spans: 3 and 4 seconds;
  • one post_ledger_entry tool span: 86 seconds, whose permission-wait child is 85 seconds and whose execution child is 1 second.
  1. Where did the time go, and what would you change?
  2. The team added a hook that runs before each tool call. With detailed tracing on, which span shows the hook's own time?
  3. The job's spans appear as a separate trace from the application that started them. What connects them?

Acceptance checks

  • Every Stage 1 answer names the class and the action, and the code catches the SDK's status-error type, branches on the status code, and logs the request ID.
  • The stop-reason function branches on stop_reason, marks truncations as incomplete, resends on pause_turn, and routes refusals to stop_details and a fallback model.
  • Each Stage 3 fix is placed correctly: integration, description, schema, or where the instructions go.
  • The trace answers use what each span covers, not guesses.

Reference solution

Stage 1

ResponseClass and action
401The key: fix the credential, and do not retry
529Transient overload: retry with backoff
402Billing: fix the payment details, and do not retry
409Conflict: resolve it, then retry
python
from anthropic import APIStatusError try: msg = client.messages.create( **request) except APIStatusError as e: rid = e.response.headers.get( "request-id") log.error("status %s id %s", e.status_code, rid) code = e.status_code if code in (401, 402, 403): alert_owner() # fix access raise if code == 409: msg = reload_and_retry() else: raise

The SDK already retries connection errors, rate limits and 5xx errors with backoff, so the handler branches on status_code and re-raises rather than looping. 401, 402 and 403 alert the owner to fix the key, billing or access; 409 resolves the conflict before its retry. anthropic.APIStatusError and status_code are the names the stop-reasons page uses in Python, and the Python SDK page reads the ID with response.headers.get("request-id"). Error bodies carry the same ID as request_id. The helper names are placeholders.

Stage 2

python
NOTE = "\n[output incomplete]" def handle(msg, history): reason = msg.stop_reason if reason == "max_tokens": return text_of(msg) + NOTE if reason == "pause_turn": history.append({ "role": "assistant", "content": msg.content}) return None # loop resends if reason == "refusal": log.info("refusal: %s", msg.stop_details) return None # try fallback return text_of(msg) # end_turn

A refusal arrives as a normal HTTP 200, so the status alone would pass it on as an answer; the documented action is to read stop_details and retry on a fallback model. A tool_use stop waits on your own tool, so your code runs lookup_customer and sends a tool_result. pause_turn is a server-tool loop pausing, so you resend the response as-is.

Stage 3

  1. Integration: send one message whose content is both tool_result blocks, before any text.
  2. The model's input: sharpen the tool descriptions so they say when to use each tool.
  3. The model's input: shrink the enum, or add input_examples showing the valid choices.
  4. Integration: move the instruction out of the tool result and keep the result to data. The page names two places: a user turn after the tool_result block, or, on supported models, a mid-conversation system message.

Stage 4

  1. 85 of the 86 seconds were the permission wait, not the tool. Pre-approve the call for the nightly job, or run it in a mode suited to unattended work.
  2. The span that wraps each hook execution; it needs detailed beta tracing.
  3. Start the agent run while an application span is active. The SDK propagates W3C trace context, so the run appears inside the application's trace.

Sources

Claude API errors; handling stop reasons; troubleshooting tool use; the Python SDK; Agent SDK observability.

Ready to study Claude Certified Developer - Foundations (CCDV-F)?

Practice tests, flashcards, and all study notes — free, no sign-up needed.

Start Studying — Free