Study Guide1,222 words

Security and Safety — study note

Security and Safety — study note

This note covers Domain 7 of the Claude Certified Developer – Foundations exam. It has four skills:

  • AI application security;
  • guardrails and safe deployment;
  • Claude hooks;
  • identity, secrets and key management.

It summarizes what each topic teaches and the Anthropic pages behind it. On purpose, it teaches mechanisms rather than numbers. Version numbers, token lifetimes and product defaults change often, so look them up on the cited pages when you need them.

Two threat models

Threat modelThe adversaryFirst defences
Jailbreak or direct injectionThe user of your applicationScreen input; tell Claude how to refuse
Indirect prompt injectionWhoever wrote content Claude readsLabel and screen what tools return
  • Harmlessness screen. A lightweight model pre-screens user input before the main conversation. Structured outputs constrain its answer to a simple classification your code can branch on.
  • Input validation. Filter input for known injection patterns. An LLM can build a generalized screen from examples of known jailbreaking language.
  • Refusals. The system prompt states ethical and legal boundaries and exactly how to refuse. A user who triggers the same refusal again and again is told their actions violate the usage policies, and may be throttled or banned.

Handling tool output

StepWhy
Say what the content is and where it came fromClaude calibrates how much to trust it
Screen each raw tool output with a small classifierCatches instructions that try to redirect the agent
Return it only if the screen is cleanOtherwise return an error or a stripped summary

Ask the screen only whether redirecting instructions are present, not whether they would succeed. Red-team the workflow with injected documents before launch, and analyze outputs for signs of successful injection after it.

Claude Code adds its own safeguards: web fetch runs in a separate context window, and in Manual mode suspicious commands prompt even if allowlisted. Avoid piping untrusted content straight to Claude, and use VMs for scripts that call external services.

Deploying agents behind a boundary

The agent runs inside the isolation boundary; credentials and other sensitive resources stay outside it. Network controls are defence in depth: they can block exfiltration even when the agent has been manipulated.

TechnologyTrade-off
Sandbox runtimeSimple to set up; shares the host kernel
ContainersNamespaces; strength depends on setup
gVisorIntercepts syscalls in userspace
VMsOwn kernel; hypervisor quality matters
  • Hardened container. --network none leaves only a mounted proxy socket; --read-only makes the root filesystem immutable; --cap-drop ALL and a process limit close escalation and fork bombs.
  • Read-only is not secret-free. Exclude credential files such as .env and copy only the source files the agent needs.
  • Allowlists read hostnames. The sandbox runtime's proxy does not inspect encrypted traffic, so domain fronting is possible; a TLS-terminating proxy gives a stronger guarantee.

Permission rules

Rules are checked deny, then ask, then allow; the first match decides, and specificity does not change the order. An allow rule cannot carve an exception out of a deny rule.

RuleEffect
A bare tool nameRemoves the tool from Claude's context
Bash(rm *)Tool stays; matching calls blocked
A runner with a wildcardMatches any inner command
A compound commandEach subcommand must match
  • Network. Argument patterns such as a curl URL prefix are fragile. Deny curl and wget, allow WebFetch by domain, and pair it with the sandbox network allowlist.
  • Read deny rules cover the built-in file tools, recognized Bash file commands such as cat, and redirect targets. They do not cover a script that opens the file itself.

The sandbox

The operating system enforces the sandbox on the running process, for Bash, PowerShell and Monitor commands and their children. Permission rules are judged earlier, from the command string.

NeedSetting
A tool writes outside the projectallowWrite for that path
A secret inside a readable treedenyRead; the narrower path wins
A token a tool must still useMask it rather than deny it

Masking needs the proxy to terminate TLS, and Claude Code ignores mask entries in a repository's own settings files, so set them in user or managed settings.

For an organization-wide gate, set the sandbox to fail when it is unavailable and turn off unsandboxed retries, so nothing silently runs outside it.

Permission modes

GoalStarting point
Review every actionManual mode
Fewer prompts, no classifierManual plus sandbox auto-allow
Hands-off with safety checksAuto mode
Unattended in a containerSkip permissions, as non-root

A default mode sets where sessions start; people can still switch. To remove a mode for everyone, set its disable key in managed settings.

Hooks

  • Claude Code. PreToolUse runs before every tool call; permission-request hooks run only when Claude Code is about to prompt. An if condition narrows a handler further. Files added with @ run no tool call, so block them with a Read deny rule. A silent hook approves nothing. A ConfigChange hook can block settings edits, except managed policy.
  • Agent SDK. Matchers match tool names only, so check the file path inside the callback. All matching hooks run in parallel in no fixed order, so write each to stand alone. permissionDecisionReason tells the model why; systemMessage tells the user.
  • Redaction. Intercept outbound tool inputs at PreToolUse and inbound results at PostToolUse.

Keys, federation and proxies

Key typeStops working
PersonalWhen you leave the organization
Service accountWhen the account is removed
Workspace (legacy)Not when its creator leaves
  • Use a personal key for your own development and a service account key for anything shared. The full key is shown once; store it in a secrets manager.
  • Workload Identity Federation exchanges your identity provider's short-lived token for a short-lived Anthropic token: no static secret to store or rotate. A leftover API key in the environment silently wins, so unset it everywhere before deleting the key. Mint a fresh identity token for each exchange.
  • Proxy injection. For Claude API calls, point the base URL at the proxy so it sees plaintext; a plain HTTPS proxy only sees a tunnel. For other services, a custom tool or MCP server can authenticate outside the boundary, or a TLS-terminating proxy can, with its CA certificate trusted.

Sources

Ready to study Claude Certified Developer - Foundations (CCDV-F)?

Practice tests, flashcards, and all study notes — free, no sign-up needed.

Start Studying — Free