Applications and Integration — practice exercise
Applications and Integration — practice exercise
Difficulty: intermediate · Estimated duration: 75–95 minutes
You are the engineer responsible for ShelfScan, an invented internal service for a chain of hardware stores. Staff photograph shelves with a phone, and ShelfScan answers their questions about stock in a chat. You will set its requirements, plan for the model life cycle, write its request code, choose how its heavier jobs run, and then harden it: client and rate-limit handling, output and sessions, and the configuration that ships with it. Every stage can be completed without an API key and without spending anything: you read the material given here, write code or a decision, and check it against the reference solution. Stage 3 has an optional step: if you have your own API key you may run your function, which uses your own credit.
What you need: a text editor and a notes file. Python is used for the code; any language is fine for the reasoning.
This exercise covers all of Domain 2: Stages 1–4 cover requirements, life cycle and API mechanics; Stages 5–7 cover engineering foundations, application design and configuration.
| Stage | What you practise | Minutes |
|---|---|---|
| 1 | Criteria, evaluations and a starting guide | 14 |
| 2 | A life-cycle plan | 12 |
| 3 | The request and the stream | 12 |
| 4 | Images, thinking, batches, platform | 10 |
| 5 | Client, limits and delivery | 14 |
| 6 | Instructions, output, sessions, plugins | 14 |
| 7 | Instruction files, settings, pins | 10 |
Stage 1 — Requirements into criteria
Skills: CCDVF-U2.T1.LO1.S1, CCDVF-U2.T1.LO1.S2, CCDVF-U2.T1.LO1.S3 Minutes: 14
The brief from the operations director reads: “ShelfScan should count stock accurately and talk to staff politely.” The team wants to start writing the prompt today.
- Name the two things the team should produce before the first prompt, in order.
- Rewrite the brief as two success criteria: one quantitative criterion for counting and one criterion for politeness. The politeness criterion may use a qualitative scale; say what must be true of how that scale is applied.
- Which kind of evaluation fits a subjective judgement such as politeness, and what does it use to judge?
- ShelfScan is a chat. Name the edge case the criteria guide lists specifically for chat applications, and write one test case of that kind.
- How should you structure the evaluation's questions where possible, and why?
- Where can the team find a production guide to adapt for the staff chat, and which one fits best?
Stage 2 — Plan for the life cycle
Skills: CCDVF-U2.T1.LO2.S1, CCDVF-U2.T1.LO2.S2, CCDVF-U2.T1.LO2.S3 Minutes: 12
- How will the team hear that ShelfScan's model is scheduled for retirement?
- Once the model is deprecated, the director suggests waiting until the week before the retirement date to move. Give the documented reason to move sooner.
- ShelfScan calls the API with an SDK. Who sends the
anthropic-versionheader, and what promise does Anthropic make about integrations that use the API as documented? - After the move, ShelfScan's replies are noticeably shorter and more direct. Is something broken, and what should the team review?
- ShelfScan's old model had no effort parameter. What latency risk comes with leaving effort unset on the new one, and which output limit should the team revisit for requests that ran without thinking?
Stage 3 — Write the request
Skills: CCDVF-U2.T2.LO3.S1, CCDVF-U2.T2.LO3.S2 Minutes: 12
ShelfScan keeps a list history per chat. Standing rules for the assistant are in RULES.
- Write
ask(history, text)in Python. It adds the staff member's message, calls the Messages API with the rules as standing instructions, adds Claude's reply tohistoryso the next call has it, and returns the reply's text. Select the text block by itstype. - Some ShelfScan answers are long. Why might the SDK insist on streaming for requests with a large
max_tokens, and how can you still get one completeMessageback? - The Python SDK offers two styles of stream. Which, and which suits a web server handling many chats at once?
- A support engineer asks which workspace ShelfScan's key resolved to for a failed request. Where can you read that from the response?
- (Optional, uses your own API key and credit.) Run
asktwice in one chat and check that the second answer can use the first.
Stage 4 — Photos, thinking, batches and platform
Skills: CCDVF-U2.T2.LO3.S3, CCDVF-U2.T2.LO3.S4 Minutes: 10
- A staff member uploads an animated GIF of a shelf being restocked and asks what changed. What will Claude see?
- The director wants ShelfScan to report exact counts of small screws in a bin from one photo. What should the team tell them, and what does that mean for the design?
- How does Claude measure an image for token cost, and what does that suggest about very large photos?
- A nightly job re-checks every store's photos. It uses a manual thinking budget on its model. How must that budget relate to
max_tokens? - The nightly job moves to the Message Batches API. Can it keep the
speedsetting it used for fast mode? And how do you delete a batch that is still running? - One region runs ShelfScan on Google Cloud and wants request-response logging. Does turning it on give Google or Anthropic access to the content?
Stage 5 — Client, limits and delivery
Skills: CCDVF-U2.T3.LO4.S1, CCDVF-U2.T3.LO4.S2, CCDVF-U2.T3.LO4.S3, CCDVF-U2.T3.LO4.S4 Minutes: 14
ShelfScan's nightly re-check now runs thousands of requests, and the team wants it to pace itself rather than collect 429s.
- Write
call(client, **kw)in Python. It makes one Messages request through the raw-response interface, reads the remaining input tokens from the response headers, sleeps forPAUSEseconds when fewer thanLOWremain, and returns the parsed message. - The same script logs each response. Which helper turns a response into a plain dictionary? And after an SDK upgrade, a documented helper is missing: what should you check first?
- The async version of
calluses the same raw-response interface. What changes about reading the parsed body? - How precise is the remaining-tokens header, and in what format does the reset header give its time?
- A long report request returns a 504
timeout_error. What does the errors page suggest? - ShelfScan's review workflow in GitHub Actions runs on every pull request. Name three documented ways to keep its runs from costing more than planned.
- Before merging a change Claude made, what should a reviewer ask Claude for, and how can Claude help with missing tests?
Stage 6 — Instructions, output, sessions and plugins
Skills: CCDVF-U2.T4.LO5.S1, CCDVF-U2.T4.LO5.S2, CCDVF-U2.T4.LO5.S3, CCDVF-U2.T4.LO5.S4 Minutes: 14
- Store managers use Claude Code in the terminal, in VS Code and in the desktop app. How many instruction files does the team need for the three surfaces, and why?
- Managers also want a weekly slide deck of stock trends. What does the prompting guide suggest adding to that request?
- Write a JSON schema for one shelf check:
sku(string),count(integer),confidence(one ofhighorlow) andnote(string, default empty). Keep to supported features. - Where in the response does the JSON arrive, and what does the Python SDK's
messages.parse()do withoutput_format? - ShelfScan's restock agent (TypeScript, Agent SDK) keeps a chat going across several messages. How does it continue the session without tracking IDs? And when the agent later moves from a laptop to a CI worker, can the worker resume from the session ID alone?
- A manager wants a plugin from Anthropic's official marketplace. Where can they see its context cost first, and is Claude Marketplace, the website, something they add with
/plugin?
Stage 7 — Instruction files, settings and pins
Skills: CCDVF-U2.T5.LO6.S1, CCDVF-U2.T5.LO6.S2, CCDVF-U2.T5.LO6.S3, CCDVF-U2.T5.LO6.S4 Minutes: 10
- ShelfScan's repository already has a CLAUDE.md. What does
/initdo now, and what happens if two instruction files disagree? - A developer answers “Yes, and don't ask again” to a Bash prompt in the CLI. Which file records it? What does the
$schemaline in a settings file add? - A developer wants to try another model for one session without changing their default. How, in the
/modelpicker? And does the Google Cloud deployment need a different ID format from the Claude API? - ShelfScan's internal plugin depends on
stock-toolsat^2.0, and the maintainers publish2.1.0-beta.1. Will users get it? How can the maintainers preview a release tag without creating it?
Acceptance checks
- Stage 1 puts criteria and evaluations before the prompt, writes one measurable and one consistently applied qualitative criterion, names an LLM-based Likert scale, the chat edge case, automated grading and the customer support agent guide.
- Stage 2 cites email and documentation notices, the reliability of deprecated models, the SDK-sent version header and the as-documented promise, a prompt review for the more concise style, and the effort and
max_tokensreviews. - Stage 3's
asksends the wholehistorywith the rules insystem, appends both turns, selects text bytype, and the answers cover streaming for large outputs, both stream styles and the workspace response header. - Stage 4 answers each question with a mechanism, never with a price or a model name.
- Stage 5's
callreads the remaining-tokens header through the raw-response interface and returns the parsed message; the answers coverto_dict(), the installed version, awaiting on the async client, header precision and format, streaming after a 504, three cost controls, and risks and edge cases. - Stage 6 keeps one instruction file for every surface, asks for design elements in the deck prompt, writes a schema with only supported features, reads the text block, and covers
continue: true, local session files, the context cost estimate and Claude Marketplace. - Stage 7 covers
/initimprovements, conflicting files, the local file,$schema, the picker'sskey, the Google Cloud format, pre-release ranges and--dry-run.
Reference solution
Stage 1. (1) Success criteria first, then evaluations that measure performance against them. The prompt comes after both. (2) For example: “Counts match a staff recount on at least 95% of a held-out set of shelf photos” and “Replies score 4 or higher on a 1–5 politeness scale on at least 90% of test chats.” The qualitative scale is valuable only if it is applied consistently, along with quantitative measures. (3) An LLM-based Likert scale: a psychometric scale that uses an LLM to judge subjective attitudes or perceptions. (4) Poor, harmful or irrelevant user input. For example, a staff member asks ShelfScan to write a complaint about a colleague. (5) Structure them for automated grading, such as multiple choice, string match, code-graded or LLM-graded, so the set can grow without hand grading. (6) Anthropic's use-case guides, which are in-depth production guides for common use cases. The customer support agent guide fits: it covers context-aware chatbots for support interactions.
Stage 2. (1) Impacted customers are always notified by email and in the documentation. (2) Deprecated models are likely to be less reliable than active ones; moving to an active model keeps the highest level of support and reliability. (3) The SDK sends it automatically. If you use the API as documented, Anthropic generally will not break your usage. (4) Not necessarily: newer models have a more concise, direct communication style, so review the prompts against the prompting best practices. (5) Higher latency at the default effort level, so consider adjusting effort as you migrate; and review max_tokens for workloads that previously ran without thinking.
Stage 3. (1) A good answer:
import anthropic
client = anthropic.Anthropic()
def ask(history, text):
history.append(
{"role": "user",
"content": text})
r = client.messages.create(
model=MODEL,
max_tokens=1024,
system=RULES,
messages=history)
history.append(
{"role": "assistant",
"content": r.content})
return next(
b.text for b in r.content
if b.type == "text")The API keeps nothing between calls, so history must carry both turns. Selecting by type survives replies that begin with thinking blocks. (2) The SDKs require streaming for large max_tokens to avoid HTTP timeouts. They can stream internally and still return the complete Message. (3) Sync and async streams; async suits a server handling many chats at once. (4) The anthropic-workspace-id response header, which names the workspace the key resolved to. (5) The second answer should refer to the first.
Stage 4. (1) Only the first frame; animations are unsupported. (2) Counts of many small objects are approximate, not precise. Design for approximate counts, for example a range or a prompt to recount by hand. (3) In patches rather than pixels; each patch is a visual token, so very large photos cost more tokens. (4) The budget must leave room for the final response, because thinking tokens count toward the turn's max_tokens. (5) No: fast mode tunes synchronous latency, which does not apply to batch processing, so drop speed. To delete a running batch, cancel it first. (6) No. Turning on the logging service gives neither Google nor Anthropic access to the content.
Stage 5. (1) A good answer:
import time
def call(client, **kw):
raw = (client.messages
.with_raw_response
.create(**kw))
left = int(raw.headers.get(
"anthropic-ratelimit-"
"input-tokens-remaining",
"0"))
if left < LOW:
time.sleep(PAUSE)
return raw.parse()House guidance: pace before the limit instead of waiting for a 429. (2) to_dict() (or to_json()) on the response. If a helper is missing after an upgrade, the environment is probably still using an older version; print anthropic.__version__. (3) On the async client the raw response is an AsyncAPIResponse, so .parse() must be awaited. (4) Rounded to the nearest thousand; the reset time is in RFC 3339 format. (5) The request timed out while processing; consider the streaming Messages API for long-running requests. (6) Any three of: --max-turns in claude_args, workflow-level timeouts, GitHub's concurrency controls, specific requests, issue templates, a concise CLAUDE.md. (7) Ask Claude to highlight potential risks or considerations, and to identify edge cases you might have missed; its tests follow the style of your existing test files.
Stage 6. (1) One: the agentic loop, tools and capabilities are the same on every surface, so one project CLAUDE.md serves them all. (2) Thoughtful design elements, visual hierarchy, and engaging animations where appropriate. (3) A good answer:
{
"type": "object",
"properties": {
"sku": {"type": "string"},
"count": {"type": "integer"},
"confidence": {
"type": "string",
"enum": ["high", "low"]},
"note": {
"type": "string",
"default": ""}
},
"required": ["sku", "count",
"confidence", "note"],
"additionalProperties": false
}(4) In the response's text content block. The Python SDK accepts output_format as a convenience and translates it to output_config.format. (5) Pass continue: true on each later query() call. No: session files are local to the machine that created them, so the ID alone is not enough on the worker. (6) In the Marketplaces tab of /plugin, where official-marketplace plugins show a context cost estimate. No: a plugin marketplace isn't Claude Marketplace.
Stage 7. (1) /init suggests improvements rather than overwriting the file. If two files give different guidance, Claude may pick either, so remove the conflict. (2) Only the local settings file. The $schema line gives autocomplete and inline validation in editors that support JSON schema. (3) Press s in the /model picker to switch without saving a default. No: on Google Cloud the format matches the Claude API. (4) No: a range doesn't match pre-release versions unless it opts in with a pre-release suffix. Pass --dry-run to claude plugin tag.
Sources
- Define success criteria and build evaluations — https://platform.claude.com/docs/en/test-and-evaluate/develop-tests
- Guides to common use cases — https://platform.claude.com/docs/en/about-claude/use-case-guides/overview
- Model deprecations — https://platform.claude.com/docs/en/about-claude/model-deprecations
- Versions — https://platform.claude.com/docs/en/api/versioning
- Model IDs and versioning — https://platform.claude.com/docs/en/about-claude/models/model-ids-and-versions
- Migration guide — https://platform.claude.com/docs/en/models/sonnet-5/migration-guide
- Using the Messages API — https://platform.claude.com/docs/en/build-with-claude/working-with-messages
- API overview — https://platform.claude.com/docs/en/api/overview
- Streaming messages — https://platform.claude.com/docs/en/build-with-claude/streaming
- Vision — https://platform.claude.com/docs/en/build-with-claude/vision
- Extended thinking — https://platform.claude.com/docs/en/build-with-claude/extended-thinking
- Batch processing — https://platform.claude.com/docs/en/build-with-claude/batch-processing
- Claude on Google Cloud — https://platform.claude.com/docs/en/build-with-claude/claude-on-vertex-ai
- Python SDK — https://platform.claude.com/docs/en/cli-sdks-libraries/sdks/python
- API errors — https://platform.claude.com/docs/en/api/errors
- Rate limits — https://platform.claude.com/docs/en/api/rate-limits
- Common workflows — https://code.claude.com/docs/en/common-workflows
- Claude Code GitHub Actions — https://code.claude.com/docs/en/github-actions
- Prompting best practices — https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/claude-prompting-best-practices
- How Claude Code works — https://code.claude.com/docs/en/how-claude-code-works
- Structured outputs — https://platform.claude.com/docs/en/build-with-claude/structured-outputs
- Work with sessions — https://code.claude.com/docs/en/agent-sdk/sessions
- Plugins overview — https://code.claude.com/docs/en/plugins/overview
- Memory — https://code.claude.com/docs/en/memory
- Settings — https://code.claude.com/docs/en/settings
- Plugin dependencies — https://code.claude.com/docs/en/plugins/dependencies