Study Guide1,124 words

Model Selection and Optimization — study note

Model Selection and Optimization — study note

This note covers Domain 5 of the Claude Certified Developer – Foundations exam. It has four skills:

  • LLM fundamentals;
  • technical fundamentals;
  • model selection and tradeoffs;
  • cost and token management.

It summarizes what each topic teaches and the Anthropic pages behind it. On purpose, it teaches mechanisms rather than numbers. Model names, prices, context-window sizes and rate limits change often, so look them up on the model and pricing pages when you need them.

What fills the context window

The context window is the model's working memory for a request, separate from the data it was trained on.

Counts toward the windowNote
System prompt and tool definitionsSent with every request
Every message, including tool results, images and documentsEarlier turns stay in full; some models strip earlier thinking blocks
The output, including thinkingThinking comes out of max_tokens
  • More context is not automatically better. Accuracy and recall degrade as the token count grows, so curate what each request needs to see.
  • Read usage. Every response reports what it consumed. With caching on, input_tokens, cache_read_input_tokens and cache_creation_input_tokens all count toward the window.
  • Thinking in tool loops. When you post tool results, return the thinking block unmodified, signature included. A modified block is an error.

Thinking, effort and fast mode

LeverWhat it doesWatch for
Effort (on models that support it)Trades thoroughness for tokens on one modelA behavioural signal, not a strict budget; applies to every output token, tool calls included
Manual thinking budgetPredictable latency and thinking cost, where supportedDeprecated or rejected on newer models; where both modes exist, use adaptive thinking with effort. A target, not a cap; max_tokens is the hard ceiling
Fast modeSame model, faster output tokens per secondNot time to first token; premium pricing; check the fast mode page for its preview status, platforms and access, which change

After a model change, run a fresh effort sweep on your evals rather than reusing old settings.

Making variable output consistent

NeedTechnique
Guaranteed JSON schemaStructured outputs (on supported models)
Stable format and toneExamples of the desired output plus an exact format
Stable factsRetrieval from a fixed information set
Consistency across a scaled workflowSmaller, consistent subtasks

A few well-crafted examples (few-shot or multishot) improve accuracy and consistency. The guide suggests three to five. For reasoning tasks, thinking tags inside the examples show the reasoning pattern.

Integrating: SDKs, streaming and transports

  • SDKs over raw HTTP. The API is RESTful. The SDKs send the authentication, version and content-type headers for you, and add type safety, retries, streaming and error handling. You pass anthropic-workspace-id yourself when your key needs it.
  • Errors and retries (Python SDK). A 4xx or 5xx response raises a subclass of APIError. Connection errors, 408, 409, 429 and 5xx are retried by default. max_retries configures that.
  • Layers. With the client SDKs you send each request and handle each response. The Agent SDK and Claude Code provide the loop, tool execution and runtime.
Stream eventHandling
message_start … message_stopA series of content blocks sits between them
Tool-input deltasPartial JSON: accumulate, then parse when the block stops
ping, errors, unknown typesExpect all three; never crash on an unknown type

Usage in message_delta is cumulative. For large max_tokens, stream, or use the SDK's final-message helper, to avoid HTTP timeouts.

Remote MCP server in Claude Code…Transport
Only answers requests, or needs OAuthHTTP (recommended)
Must push events unpromptedWebSocket

These transport rules come from Claude Code's MCP page: in Claude Code, HTTP supports OAuth and the add-transport flag, while WebSocket supports neither. They describe Claude Code's client, not the protocol itself.

Choosing and changing models

  • Criteria. Balance capabilities, speed and cost. Tuning effort is often a better lever than switching models.
  • Starting point.
    • Capability-first: the strongest start, then optimize down.
    • Efficiency-first: start small and upgrade only for specific gaps.
  • Evidence. A good evaluation set is the most important step. Test with your actual prompts and data. Every model ID is a pinned snapshot.
  • Cost. Compare on cost per completed task. Price the hardest tenth of your tasks, because a failed task bills its tokens, then the retry, then whatever the failure costs downstream.
  • Multi-model. Sweep effort first; most workloads end there. If a gap remains, price the stronger model alone at low effort: that is the baseline any multi-model design must beat. An orchestrator fits work that splits into independent pieces; an advisor fits serial work that is hard in spots. Uniform work, or one dependent chain, usually wants one well-tuned model.
  • Upgrades. Read the breaking changes, select content blocks by type, audit prompts written for the old model, and re-sweep effort.

Controlling the bill

MechanismKey rule
Prompt cachingAutomatic for growing conversations; explicit breakpoints for fine control; reads bill at the cache-read rate; changing effort or thinking invalidates it onward (on some models, the tools and system prompt too), but setting effort to its default explicitly does not; editing tool definitions invalidates everything
Token countingAn estimate; recount per target model; the endpoint errors on most server tools and the MCP connector, so read usage for those
Usage tracking (Agent SDK)Deduplicate parallel tool calls by message ID; read the latest result for a session total
BatchesFor work no one waits on; match results by custom_id, since order is not guaranteed

Sources

Ready to study Claude Certified Developer - Foundations (CCDV-F)?

Practice tests, flashcards, and all study notes — free, no sign-up needed.

Start Studying — Free