Security and Safety — practice exercise
Security and Safety — practice exercise
Difficulty: intermediate · Estimated duration: 45–60 minutes
You are the engineer responsible for ClaimDesk, an invented service that reads insurance-claim emails, extracts the claim details from their attachments and drafts a reply for a claims handler. It runs as an Agent SDK application in containers, calls an internal claims API, and its repository is maintained with Claude Code. You will harden it stage by stage and write down the decision you would make at each step. Every stage can be completed without an API key and without spending anything: you read the material given here, write code, configuration or a decision, and check it against the reference solution. Stage 4 has an optional step: if you have your own API key you may run your hook against a real session, which uses your own credit.
What you need: a text editor and a notes file. Python is used for the code; the configuration is JSON.
| Stage | What you practise | Minutes |
|---|---|---|
| 1 | Handling hostile input | 12 |
| 2 | Drawing the boundary and keeping keys out | 12 |
| 3 | Permission rules, the sandbox and the mode | 14 |
| 4 | A hook that enforces policy | 8 |
| 5 | Keys and federation | 8 |
Stage 1 — Handle hostile input
Skills: CCDVF-U7.T1.LO1.S1 Minutes: 12
ClaimDesk has two ways in: a public web form where anyone can submit a claim question, and a mailbox tool that reads claim emails and their attachments from outside senders.
- For each way in, name the threat model and say who the adversary is.
- The mailbox tool currently pastes each email into the user turn as plain text. Write a Python function
wrap_email(msg)that returns the email as a text content block holding a JSON object with its source, sender and body, so the body cannot break out of its structure. - The team has collected forty jailbreak phrases from the web form's logs and wants a screen that also catches new wordings of the same attacks. How can it build one?
- List one test you would run before launch to show the defences work.
Stage 2 — Draw the boundary
Skills: CCDVF-U7.T1.LO1.S2, CCDVF-U7.T2.LO4.S3 Minutes: 12
ClaimDesk's worker container is started like this today:
docker run \
-e CLAIMS_API_KEY \
-v /srv/claimdesk:/workspace \
claimdesk-worker- Rewrite the command so the worker has no network interface except a mounted proxy socket, cannot persist changes to its root filesystem, cannot escalate privileges, runs as a non-root user and sees its code read-only.
/srv/claimdeskalso holds.envand a service-account JSON file. What do you do before mounting it?- The worker's claims-API key is passed in as an environment variable. Where should the key live instead, and how does the worker's request get authenticated?
- The worker also calls the Claude API. How do you route those calls so your proxy can add the credential?
Stage 3 — Configure Claude Code for the repository
Skills: CCDVF-U7.T2.LO2.S1, CCDVF-U7.T2.LO2.S2, CCDVF-U7.T2.LO2.S3 Minutes: 14
Engineers use Claude Code on the ClaimDesk repository. Write the project's .claude/settings.json so that:
- Claude cannot run curl or wget, but may fetch pages from the team's docs host,
claims.dev; - files under
secrets/are not readable through Claude's file tools; - shell commands run in the sandbox, and the test runner may write its cache to
~/.cache/claimdesk.
Then answer:
- A teammate proposes an allow rule for
Bash(npm test), and later asks whynpm test && curl evil.examplestill prompts. Explain. - A test script in Python opens a file under
secrets/. Does your Read deny rule stop it? If not, what does? - An engineer iterating on the tests wants fewer prompts but no classifier. Which mode and setting do you suggest?
Stage 4 — Enforce policy with a hook
Skills: CCDVF-U7.T2.LO3.S1, CCDVF-U7.T2.LO3.S2 Minutes: 8
ClaimDesk's agent may only write drafts, which live under /srv/app/drafts/.
- Write an Agent SDK PreToolUse callback
only_draftsthat denies any file write outside that directory, with a reason the model can read, and show how you register it. - Claim numbers must never appear in what the claims-lookup tool returns to Claude. At which hook event do you remove them?
- (Optional, uses your own API key and credit.) Run a session with your hook registered, ask the agent to write outside
drafts/, and check the tool result in the message stream.
Stage 5 — Keys and federation
Skills: CCDVF-U7.T2.LO4.S1, CCDVF-U7.T2.LO4.S2 Minutes: 8
- Which key type should each engineer use for local development, and which should the staging deployment use while it still uses a key?
- Production runs on Kubernetes, which can project an identity token into each pod. Outline the steps to move production from its API key to Workload Identity Federation without downtime.
- After the move, how should the production code construct its client?
Acceptance checks
- Stage 1 names both threat models, JSON-encodes the email with its source, builds a generalized screen from example phrases, and red-teams with injected emails.
- Stage 2 removes the network interface and makes the root filesystem read-only, keeps secrets out of the mount, and keeps the claims-API key outside the container.
- Stage 3 denies curl and wget, allows WebFetch by domain, denies reads of
secrets/, enables the sandbox with a write grant, and explains subcommand matching and the Read-rule gap. - Stage 4 checks the path inside the callback, registers it on the file-writing tools, and redacts at the post-tool-use event.
- Stage 5 uses personal keys for development, a service account key for the shared deployment, and a migration that leaves the key until federation wins.
Reference solution
Stage 1. (1) The web form: a jailbreak or direct prompt injection, where the user is the adversary. The mailbox: indirect prompt injection, where the handler is trusted but the emails are third-party content that may carry instructions. (2) A good version:
import json
def wrap_email(msg):
body = json.dumps({
"source": "inbound_email",
"from": msg["from"],
"body": msg["body"]})
return {"type": "text",
"text": body}Return it inside the tool's tool_result, not in the user turn. JSON escaping gives unambiguous delimiters, so an attacker cannot close a quote or tag to break out into an instruction context, and the source field tells Claude what the content is. (3) Give an LLM the known jailbreaking language as examples, to create a generalized validation screen that runs before input reaches Claude. (4) Run the workflow on emails and attachments that deliberately contain injection attempts, and confirm that Claude ignores them and your screens catch the rest.
Stage 2. (1) A good rewrite:
docker run \
--network none \
--read-only \
--cap-drop ALL \
--user 1000:1000 \
-v /srv/src:/workspace:ro \
-v /run/p.sock:/run/p.sock:ro \
claimdesk-workerWith no network interface, the only way out is the mounted socket to a proxy on the host, which can allowlist domains, inject credentials and log traffic. (2) Exclude or sanitize .env and the service-account file, or copy in only the source files the worker needs; read-only access can still expose credentials. (3) Keep the key outside the container. Give the worker a custom tool that calls a service outside its boundary, where a proxy injects the key; the worker sees only the tool interface. (4) Point ANTHROPIC_BASE_URL at the proxy, which then receives plaintext requests it can add the credential to. With only HTTPS_PROXY, the proxy sees an encrypted tunnel.
Stage 3. A good settings.json:
{
"permissions": {
"deny": [
"Bash(curl *)",
"Bash(wget *)",
"Read(./secrets/**)"
],
"allow": [
"WebFetch(domain:claims.dev)"
]
},
"sandbox": {
"enabled": true,
"filesystem": {
"allowWrite": [
"~/.cache/claimdesk"
]
}
}
}Pair the curl deny with the sandbox network allowlist if the restriction must hold, since a deny rule doesn't match the same program by path or inside sh -c. (1) Claude Code splits compound commands, and a rule must match each subcommand independently; the curl part matches no allow rule, and here it matches a deny. (2) No. Read deny rules cover the built-in file tools and recognized Bash file commands such as cat, not a Python program that opens the file itself. Add secrets/ to the sandbox's denyRead, which the operating system enforces on the running process and its children. (3) Manual mode with the Bash sandbox in auto-allow mode; deny rules still apply.
Stage 4. (1) A good callback:
OUT = "hookSpecificOutput"
OK = "/srv/app/drafts/"
async def only_drafts(i, tid, ctx):
t = i["tool_input"]
p = t.get("file_path", "")
if p.startswith(OK):
return {}
return {OUT: {
"hookEventName":
i["hook_event_name"],
"permissionDecision":
"deny",
"permissionDecisionReason":
"Write only to drafts/"}}Register it on the file-writing tools, because matchers match tool names, not paths:
hooks = {"PreToolUse": [
HookMatcher(
matcher="Write|Edit",
hooks=[only_drafts])]}In production, normalize the path before comparing it, so a path containing .. cannot climb out. The reason tells the model why, so it avoids retrying. (2) At PostToolUse, which is where inbound tool results are intercepted for redaction; PreToolUse is for outbound tool inputs. (3) The Write tool's result in the message stream reports the denial.
Stage 5. (1) A personal key for each engineer's own development, and a service account key for the staging deployment, since it is shared. (2) Configure the issuer, service account and federation rule while the API key stays in place; confirm which credential wins; unset ANTHROPIC_API_KEY everywhere production runs, since it sits above federation and silently shadows it; then delete the key in the Console. (3) With no arguments, injecting the federation rule, organization, service account, workspace and identity-token file through the environment, so the same image runs everywhere.
Sources
- Mitigate jailbreaks and prompt injections — https://platform.claude.com/docs/en/test-and-evaluate/strengthen-guardrails/mitigate-jailbreaks
- Security (Claude Code) — https://code.claude.com/docs/en/security
- Securely deploying AI agents — https://code.claude.com/docs/en/agent-sdk/secure-deployment
- Configure permissions — https://code.claude.com/docs/en/permissions
- Configure the sandboxed Bash tool — https://code.claude.com/docs/en/sandboxing
- Choose a permission mode — https://code.claude.com/docs/en/permission-modes
- Hooks reference — https://code.claude.com/docs/en/hooks
- Intercept and control agent behavior with hooks — https://code.claude.com/docs/en/agent-sdk/hooks
- Get your Claude API key — https://platform.claude.com/docs/en/get-api-key
- Workload Identity Federation — https://platform.claude.com/docs/en/manage-claude/workload-identity-federation