Security and Safety — study note
Security and Safety — study note
This note covers Domain 7 of the Claude Certified Developer – Foundations exam. It has four skills:
- AI application security;
- guardrails and safe deployment;
- Claude hooks;
- identity, secrets and key management.
It summarizes what each topic teaches and the Anthropic pages behind it. On purpose, it teaches mechanisms rather than numbers. Version numbers, token lifetimes and product defaults change often, so look them up on the cited pages when you need them.
Two threat models
| Threat model | The adversary | First defences |
|---|---|---|
| Jailbreak or direct injection | The user of your application | Screen input; tell Claude how to refuse |
| Indirect prompt injection | Whoever wrote content Claude reads | Label and screen what tools return |
- Harmlessness screen. A lightweight model pre-screens user input before the main conversation. Structured outputs constrain its answer to a simple classification your code can branch on.
- Input validation. Filter input for known injection patterns. An LLM can build a generalized screen from examples of known jailbreaking language.
- Refusals. The system prompt states ethical and legal boundaries and exactly how to refuse. A user who triggers the same refusal again and again is told their actions violate the usage policies, and may be throttled or banned.
Handling tool output
| Step | Why |
|---|---|
| Say what the content is and where it came from | Claude calibrates how much to trust it |
| Screen each raw tool output with a small classifier | Catches instructions that try to redirect the agent |
| Return it only if the screen is clean | Otherwise return an error or a stripped summary |
Ask the screen only whether redirecting instructions are present, not whether they would succeed. Red-team the workflow with injected documents before launch, and analyze outputs for signs of successful injection after it.
Claude Code adds its own safeguards: web fetch runs in a separate context window, and in Manual mode suspicious commands prompt even if allowlisted. Avoid piping untrusted content straight to Claude, and use VMs for scripts that call external services.
Deploying agents behind a boundary
The agent runs inside the isolation boundary; credentials and other sensitive resources stay outside it. Network controls are defence in depth: they can block exfiltration even when the agent has been manipulated.
| Technology | Trade-off |
|---|---|
| Sandbox runtime | Simple to set up; shares the host kernel |
| Containers | Namespaces; strength depends on setup |
| gVisor | Intercepts syscalls in userspace |
| VMs | Own kernel; hypervisor quality matters |
- Hardened container.
--network noneleaves only a mounted proxy socket;--read-onlymakes the root filesystem immutable;--cap-drop ALLand a process limit close escalation and fork bombs. - Read-only is not secret-free. Exclude credential files such as
.envand copy only the source files the agent needs. - Allowlists read hostnames. The sandbox runtime's proxy does not inspect encrypted traffic, so domain fronting is possible; a TLS-terminating proxy gives a stronger guarantee.
Permission rules
Rules are checked deny, then ask, then allow; the first match decides, and specificity does not change the order. An allow rule cannot carve an exception out of a deny rule.
| Rule | Effect |
|---|---|
| A bare tool name | Removes the tool from Claude's context |
Bash(rm *) | Tool stays; matching calls blocked |
| A runner with a wildcard | Matches any inner command |
| A compound command | Each subcommand must match |
- Network. Argument patterns such as a curl URL prefix are fragile. Deny curl and wget, allow WebFetch by domain, and pair it with the sandbox network allowlist.
- Read deny rules cover the built-in file tools, recognized Bash file commands such as
cat, and redirect targets. They do not cover a script that opens the file itself.
The sandbox
The operating system enforces the sandbox on the running process, for Bash, PowerShell and Monitor commands and their children. Permission rules are judged earlier, from the command string.
| Need | Setting |
|---|---|
| A tool writes outside the project | allowWrite for that path |
| A secret inside a readable tree | denyRead; the narrower path wins |
| A token a tool must still use | Mask it rather than deny it |
Masking needs the proxy to terminate TLS, and Claude Code ignores mask entries in a repository's own settings files, so set them in user or managed settings.
For an organization-wide gate, set the sandbox to fail when it is unavailable and turn off unsandboxed retries, so nothing silently runs outside it.
Permission modes
| Goal | Starting point |
|---|---|
| Review every action | Manual mode |
| Fewer prompts, no classifier | Manual plus sandbox auto-allow |
| Hands-off with safety checks | Auto mode |
| Unattended in a container | Skip permissions, as non-root |
A default mode sets where sessions start; people can still switch. To remove a mode for everyone, set its disable key in managed settings.
Hooks
- Claude Code.
PreToolUseruns before every tool call; permission-request hooks run only when Claude Code is about to prompt. Anifcondition narrows a handler further. Files added with@run no tool call, so block them with a Read deny rule. A silent hook approves nothing. AConfigChangehook can block settings edits, except managed policy. - Agent SDK. Matchers match tool names only, so check the file path inside the callback. All matching hooks run in parallel in no fixed order, so write each to stand alone.
permissionDecisionReasontells the model why;systemMessagetells the user. - Redaction. Intercept outbound tool inputs at
PreToolUseand inbound results atPostToolUse.
Keys, federation and proxies
| Key type | Stops working |
|---|---|
| Personal | When you leave the organization |
| Service account | When the account is removed |
| Workspace (legacy) | Not when its creator leaves |
- Use a personal key for your own development and a service account key for anything shared. The full key is shown once; store it in a secrets manager.
- Workload Identity Federation exchanges your identity provider's short-lived token for a short-lived Anthropic token: no static secret to store or rotate. A leftover API key in the environment silently wins, so unset it everywhere before deleting the key. Mint a fresh identity token for each exchange.
- Proxy injection. For Claude API calls, point the base URL at the proxy so it sees plaintext; a plain HTTPS proxy only sees a tunnel. For other services, a custom tool or MCP server can authenticate outside the boundary, or a TLS-terminating proxy can, with its CA certificate trusted.
Sources
- Mitigate jailbreaks and prompt injections — https://platform.claude.com/docs/en/test-and-evaluate/strengthen-guardrails/mitigate-jailbreaks
- Security (Claude Code) — https://code.claude.com/docs/en/security
- Securely deploying AI agents — https://code.claude.com/docs/en/agent-sdk/secure-deployment
- Configure permissions — https://code.claude.com/docs/en/permissions
- Configure the sandboxed Bash tool — https://code.claude.com/docs/en/sandboxing
- Choose a permission mode — https://code.claude.com/docs/en/permission-modes
- Hooks reference — https://code.claude.com/docs/en/hooks
- Intercept and control agent behavior with hooks — https://code.claude.com/docs/en/agent-sdk/hooks
- Get your Claude API key — https://platform.claude.com/docs/en/get-api-key
- Workload Identity Federation — https://platform.claude.com/docs/en/manage-claude/workload-identity-federation