Hands-on Lab1,534 words

Model Selection and Optimization — practice exercise

Model Selection and Optimization — practice exercise

Difficulty: intermediate · Estimated duration: 45–60 minutes

You are the engineer responsible for TriageDesk, an invented internal service that reads incoming support tickets, tags each one and drafts a first reply. It is high-volume and latency-sensitive: thousands of tickets a day, and agents wait for each tag. You will review its design and code and write down the decision you would make at each step. Every stage can be completed without an API key and without spending anything: you read the material given here, write code or a decision, and check it against the reference solution. Stage 2 has an optional step: if you have your own API key you may run your rewritten client, which uses your own credit.

What you need: a text editor and a notes file. Python is used for the code; any language is fine for the reasoning.

StageWhat you practiseMinutes
1Reading cache usage, sizing effort, checking examples12
2Rewriting a hand-rolled client and streaming the reply14
3Choosing a transport and keeping a slow tool alive8
4A model and cost plan16

Stage 1 — Size the request

Skills: CCDVF-U5.T1.LO1.S1, CCDVF-U5.T1.LO1.S2, CCDVF-U5.T1.LO1.S3 Minutes: 12

TriageDesk logged this usage for one request (the numbers are invented for the exercise):

  • input_tokens: 3,100
  • cache_read_input_tokens: 41,000
  • cache_creation_input_tokens: 0
  • output_tokens: 2,400

The request carried a system prompt, 22 tool definitions, a ticket and its thread history. Thinking is enabled and the model supports the effort parameter. Most tickets are simple password resets.

  1. From the table, say how many input tokens were read from the cache and how many were written to it, and whether this request hit the cache.
  2. Choose one change for the simple tickets and justify it in one sentence: lower effort, or a larger thinking budget.
  3. The tagging prompt carries four example tags, all taken from password-reset tickets, and the team is unsure they are varied enough. What can you ask Claude to do with them?

Stage 2 — Harden the client

Skills: CCDVF-U5.T2.LO2.S1, CCDVF-U5.T2.LO2.S2 Minutes: 14

TriageDesk's current client, simplified:

python
import requests def draft(ticket): r = requests.post( API_URL, json=body(ticket), headers={"x-api-key": KEY}) return r.json()
  1. The client already sends the API key. Name three other things it leaves to chance that a client SDK handles for you.
  2. Rewrite draft to use the Python SDK. Its requests use a large max_tokens, and the agent UI does not show the reply until it is complete, so avoid holding a long idle connection. Return the reply's text and raise a clear error when there is none.
  3. Agents want a short view of Claude's reasoning while a draft streams, not the full chain of thought. Which thinking display setting gives that?
  4. The stream to the agent UI sometimes drops mid-reply, after a tool call and some text. Can you resume, and from where?
  5. (Optional, uses your own API key and credit.) Run your rewritten draft once against the API and compare the usage field with your Stage 1 reasoning.

Stage 3 — Pick the transport

Skills: CCDVF-U5.T2.LO2.S3 Minutes: 8

  1. TriageDesk's engineers keep a script that reads log files on their own machines, and want Claude Code to use it as an MCP server. Which transport fits, and why?
  2. The knowledge-search tool, on a remote MCP server connected to Claude Code, can work for minutes without sending anything, and its calls are aborted with an error before their overall time limit. What should the server do?

Stage 4 — Plan the model and the bill

Skills: CCDVF-U5.T3.LO3.S1, CCDVF-U5.T3.LO3.S2, CCDVF-U5.T3.LO3.S3, CCDVF-U5.T3.LO4.S1, CCDVF-U5.T3.LO4.S2, CCDVF-U5.T3.LO4.S3 Minutes: 16

Write a one-page plan that answers each question in two or three sentences:

  1. Starting model. TriageDesk is high-volume and latency-sensitive. Which starting approach fits, and why?
  2. Effort and cost. A tag can be checked against the ticket's final resolution. How do you set effort to keep cost down, and why might a pricier model still cost less per ticket?
  3. Upgrade. List three checks for moving TriageDesk to a newer model.
  4. Caching. The system prompt and tool definitions are cached. What must stay unchanged between requests for the cache to hold?
  5. Counting. Some tickets attach screenshots by URL. How do you count their tokens before sending?
  6. Tracking and batching. A nightly job re-tags last week's tickets. How do you run it, and how do you avoid double-counting cost in your Agent SDK worker?

Acceptance checks

  • Stage 1 reads the cache fields correctly, picks lower effort for simple tickets, and asks Claude to review the examples for relevance and diversity.
  • Stage 2 uses the SDK, streams internally and returns the complete message, selects content by type, picks the summarized display, and resumes from the most recent text block.
  • Stage 3 picks stdio for the local script and progress notifications for the slow tool, with the reason for each.
  • Stage 4 answers all six questions with a mechanism, never with a price or a model name.

Reference solution

Stage 1. (1) cache_read_input_tokens is the number of tokens retrieved from the cache for this request, so 41,000 input tokens came from the cache. cache_creation_input_tokens is the number written when creating a new entry, so nothing was written. The request hit the cache for its prefix, and only 3,100 input tokens were processed uncached. (2) Lower effort, since this model supports it. Effort applies to every output token, including tool calls, and it is the lever for trading thoroughness against tokens on a single model. A larger budget spends more. (3) Ask Claude to evaluate the examples for relevance and diversity, or to generate additional ones from the initial set, so the examples cover more than password resets.

Stage 2. (1) Any three of: the version and content-type headers; retries; timeouts; error handling; streaming helpers. (2) A good rewrite:

python
import anthropic # reads the key from the environment client = anthropic.Anthropic() def draft(ticket): with client.messages.stream( **body(ticket)) as s: msg = s.get_final_message() text = next( (b.text for b in msg.content if b.type == "text"), None) if text is None: raise ValueError( msg.stop_reason) return text

Some networks drop idle connections, so a large max_tokens without streaming can fail. The SDK streams internally and returns the complete message, which is what the UI needs. A reply can begin with thinking blocks, so the rewrite selects content blocks by type rather than by position. (3) display: "summarized", which streams a condensed summary of Claude's reasoning rather than the full chain of thought. (4) Yes, partly. Tool use and extended thinking blocks cannot be partially recovered; resume streaming from the most recent text block. (5) The response's usage field reports what the request consumed.

Stage 3. (1) Stdio. Stdio servers run as local processes on your machine and suit tools that need direct system access or custom scripts. (2) Send progress notifications while it works. A tool call that sends no response and no progress notification for the idle window aborts with an error instead of waiting for the wall-clock limit.

Stage 4. (1) Efficiency-first: start with a faster, more cost-effective model. The selection guide lists applications with tight latency requirements among those that approach is best for. (2) Run every ticket at a low effort setting and re-run only the failures at a higher one, which needs a checker that does not pass bad tags. Compare models on cost per completed task: a more capable model can finish with less work, with fewer turns and less backtracking. (3) Read the migration guide's breaking changes; select content blocks by type, since a reply can begin with thinking blocks; and re-baseline cost on your typical workload before production. (4) Keep the tool definitions unchanged, because editing them invalidates the entire cache. If you use a task budget, set it once on the first request, because changing it partway invalidates any cached prefix that contains it. Keep the effort level and thinking configuration fixed too, since changing them between requests invalidates the cache from that point onward. (5) Send the screenshots as base64 to the counting endpoint; it does not accept images given by URL. The count is an estimate. (6) Send it through the batch API, since no one is waiting on it, and keep the interactive path for live tickets. In the Agent SDK, deduplicate usage by message ID, because messages from parallel tool calls in one turn share an ID.

Sources

Ready to study Claude Certified Developer - Foundations (CCDV-F)?

Practice tests, flashcards, and all study notes — free, no sign-up needed.

Start Studying — Free