Model Selection and Optimization — study note
Model Selection and Optimization — study note
This note covers Domain 5 of the Claude Certified Developer – Foundations exam. It has four skills:
- LLM fundamentals;
- technical fundamentals;
- model selection and tradeoffs;
- cost and token management.
It summarizes what each topic teaches and the Anthropic pages behind it. On purpose, it teaches mechanisms rather than numbers. Model names, prices, context-window sizes and rate limits change often, so look them up on the model and pricing pages when you need them.
What fills the context window
The context window is the model's working memory for a request, separate from the data it was trained on.
| Counts toward the window | Note |
|---|---|
| System prompt and tool definitions | Sent with every request |
| Every message, including tool results, images and documents | Earlier turns stay in full; some models strip earlier thinking blocks |
| The output, including thinking | Thinking comes out of max_tokens |
- More context is not automatically better. Accuracy and recall degrade as the token count grows, so curate what each request needs to see.
- Read
usage. Every response reports what it consumed. With caching on,input_tokens,cache_read_input_tokensandcache_creation_input_tokensall count toward the window. - Thinking in tool loops. When you post tool results, return the thinking block unmodified, signature included. A modified block is an error.
Thinking, effort and fast mode
| Lever | What it does | Watch for |
|---|---|---|
| Effort (on models that support it) | Trades thoroughness for tokens on one model | A behavioural signal, not a strict budget; applies to every output token, tool calls included |
| Manual thinking budget | Predictable latency and thinking cost, where supported | Deprecated or rejected on newer models; where both modes exist, use adaptive thinking with effort. A target, not a cap; max_tokens is the hard ceiling |
| Fast mode | Same model, faster output tokens per second | Not time to first token; premium pricing; check the fast mode page for its preview status, platforms and access, which change |
After a model change, run a fresh effort sweep on your evals rather than reusing old settings.
Making variable output consistent
| Need | Technique |
|---|---|
| Guaranteed JSON schema | Structured outputs (on supported models) |
| Stable format and tone | Examples of the desired output plus an exact format |
| Stable facts | Retrieval from a fixed information set |
| Consistency across a scaled workflow | Smaller, consistent subtasks |
A few well-crafted examples (few-shot or multishot) improve accuracy and consistency. The guide suggests three to five. For reasoning tasks, thinking tags inside the examples show the reasoning pattern.
Integrating: SDKs, streaming and transports
- SDKs over raw HTTP. The API is RESTful. The SDKs send the authentication, version and content-type headers for you, and add type safety, retries, streaming and error handling. You pass
anthropic-workspace-idyourself when your key needs it. - Errors and retries (Python SDK). A 4xx or 5xx response raises a subclass of
APIError. Connection errors, 408, 409, 429 and 5xx are retried by default.max_retriesconfigures that. - Layers. With the client SDKs you send each request and handle each response. The Agent SDK and Claude Code provide the loop, tool execution and runtime.
| Stream event | Handling |
|---|---|
message_start … message_stop | A series of content blocks sits between them |
| Tool-input deltas | Partial JSON: accumulate, then parse when the block stops |
ping, errors, unknown types | Expect all three; never crash on an unknown type |
Usage in message_delta is cumulative. For large max_tokens, stream, or use the SDK's final-message helper, to avoid HTTP timeouts.
| Remote MCP server in Claude Code… | Transport |
|---|---|
| Only answers requests, or needs OAuth | HTTP (recommended) |
| Must push events unprompted | WebSocket |
These transport rules come from Claude Code's MCP page: in Claude Code, HTTP supports OAuth and the add-transport flag, while WebSocket supports neither. They describe Claude Code's client, not the protocol itself.
Choosing and changing models
- Criteria. Balance capabilities, speed and cost. Tuning effort is often a better lever than switching models.
- Starting point.
- Capability-first: the strongest start, then optimize down.
- Efficiency-first: start small and upgrade only for specific gaps.
- Evidence. A good evaluation set is the most important step. Test with your actual prompts and data. Every model ID is a pinned snapshot.
- Cost. Compare on cost per completed task. Price the hardest tenth of your tasks, because a failed task bills its tokens, then the retry, then whatever the failure costs downstream.
- Multi-model. Sweep effort first; most workloads end there. If a gap remains, price the stronger model alone at low effort: that is the baseline any multi-model design must beat. An orchestrator fits work that splits into independent pieces; an advisor fits serial work that is hard in spots. Uniform work, or one dependent chain, usually wants one well-tuned model.
- Upgrades. Read the breaking changes, select content blocks by
type, audit prompts written for the old model, and re-sweep effort.
Controlling the bill
| Mechanism | Key rule |
|---|---|
| Prompt caching | Automatic for growing conversations; explicit breakpoints for fine control; reads bill at the cache-read rate; changing effort or thinking invalidates it onward (on some models, the tools and system prompt too), but setting effort to its default explicitly does not; editing tool definitions invalidates everything |
| Token counting | An estimate; recount per target model; the endpoint errors on most server tools and the MCP connector, so read usage for those |
| Usage tracking (Agent SDK) | Deduplicate parallel tool calls by message ID; read the latest result for a session total |
| Batches | For work no one waits on; match results by custom_id, since order is not guaranteed |
Sources
- Context windows — https://platform.claude.com/docs/en/build-with-claude/context-windows
- Effort — https://platform.claude.com/docs/en/build-with-claude/effort
- Extended thinking — https://platform.claude.com/docs/en/build-with-claude/extended-thinking
- Fast mode — https://platform.claude.com/docs/en/build-with-claude/fast-mode
- Prompting best practices — https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/claude-prompting-best-practices
- Increase output consistency — https://platform.claude.com/docs/en/test-and-evaluate/strengthen-guardrails/increase-consistency
- CLI, SDKs and libraries — https://platform.claude.com/docs/en/cli-sdks-libraries/overview
- Python SDK — https://platform.claude.com/docs/en/cli-sdks-libraries/sdks/python
- API overview — https://platform.claude.com/docs/en/api/overview
- Streaming messages — https://platform.claude.com/docs/en/build-with-claude/streaming
- Connect Claude Code to tools via MCP — https://code.claude.com/docs/en/mcp
- Hosting the Agent SDK — https://code.claude.com/docs/en/agent-sdk/hosting
- Choosing the right model — https://platform.claude.com/docs/en/about-claude/models/choosing-a-model
- Models overview — https://platform.claude.com/docs/en/models/overview
- Optimizing for cost and intelligence — https://platform.claude.com/docs/en/about-claude/models/optimizing-for-cost-and-intelligence
- Migration guide (breaking changes) — https://platform.claude.com/docs/en/models/sonnet-5/migration-guide
- Prompt caching — https://platform.claude.com/docs/en/build-with-claude/prompt-caching
- Token counting — https://platform.claude.com/docs/en/build-with-claude/token-counting
- Track cost and usage — https://code.claude.com/docs/en/agent-sdk/cost-tracking
- Batch processing — https://platform.claude.com/docs/en/build-with-claude/batch-processing