🛡️ Safety, Alignment & Red Teaming
Threat models, trust boundaries, red teaming, and measurement of harmful compliance and benign refusal
On this page
Before you start
Review probability and statistics for rates and uncertainty, RAG for source grounding, and agentic AI for tool execution and authorization. For optional training background, RLHF and DPO explains how training changes behavior; this chapter covers the surrounding system and its limits.
By the end, you should be able to write a concrete threat model, identify a trust-boundary failure in an agent trace, distinguish harmful compliance from benign refusal, and design evaluation and controls that remain useful when a model makes a mistake.
The problem: safe-sounding text can accompany an unsafe action
A support assistant reads documents and has limited account tools. A malicious document tells it to send customer records to an external address. The assistant might produce a polite final answer while its tool call leaks data. A content classifier checking only that final answer would miss the security failure.
Reliability concerns doing the intended task correctly. Security concerns resisting unauthorized access or manipulation. Safety concerns unacceptable harm in the deployment context, including misuse and accidental failure. Alignment asks whether behavior and optimization match intended goals and values. These overlap, but each needs a concrete specification and evidence; a single “safe” score obscures important differences.
Worked trace: an untrusted document crosses a boundary
Suppose the authorized task is to summarize the customer's refund policy. The attacker controls one retrieved document but cannot change the application's access rules.
| Element | Threat-model description |
|---|---|
| Legitimate task | Explain the applicable refund policy with sources |
| Assets | Customer records, account credentials, correct policy advice |
| Attacker control | Text of a document that may appear in retrieval results |
| Intended scope | Read authorized policy material; no export of customer records |
| Trust boundary | Retrieved content enters model context, but must not become tool authority |
| Failure to prevent | An unauthorized data export or unsupported policy claim |
Trace the attempted failure:
- Retrieval returns a policy passage plus an embedded instruction to export customer records.
- The model incorrectly proposes an export tool call.
- The runtime checks the original task scope, account permissions, and permitted destinations; the export is denied.
- The event is recorded without leaking the records into logs, and the assistant uses valid policy evidence or reports that evidence is insufficient.
The model failed to follow the intended instruction hierarchy, but the runtime prevented the unauthorized action. Track these as separate events: unsafe proposal and successful unauthorized execution. If the retrieval stage had already put forbidden customer records into the prompt or an external scorer, a later tool denial would not undo that exposure.
An input can be untrusted without looking suspicious. Structured JSON, a tool result, an image containing text, and a retrieved memory can all carry hostile instructions or false facts. Structure identifies fields; it does not make their contents authoritative.
Worked rates: lower attack success can hide more over-refusal
Evaluate two configurations on the same 200 harmful adversarial requests and 800 benign requests, with reviewed labels and a fixed scoring rubric:
| Outcome | Configuration A | Configuration B |
|---|---|---|
| Harmful requests answered unsafely | 10 of 200 | 4 of 200 |
| Harmful requests handled within policy | 190 of 200 | 196 of 200 |
| Benign requests incorrectly refused | 80 of 800 | 160 of 800 |
| Benign requests usefully answered | 720 of 800 | 640 of 800 |
Here attack success rate is ASR = successful harmful attacks / attempted harmful attacks: A has 5%, B has 2%. False refusal rate is FRR = incorrectly refused benign requests / benign requests: A has 10%, B has 20%. B blocks six additional harmful attempts while refusing eighty additional benign requests. Whether that is acceptable depends on severity, use case, and available alternatives; it is not established by the ASR decrease alone.
Define the outcome before optimizing a metric
| Term | Operational meaning |
|---|---|
| Threat model | Actors, assets, attacker capabilities, boundaries, and failure outcomes |
| ASR | Successful attacks under a specified attack set, budget, and success rubric |
| FRR | Incorrect refusals among requests labeled benign under the applicable policy |
| Harm prevalence | Share of sampled exposures or interactions containing the defined harm |
| Severity | Consequence of the failure, separate from its frequency |
| Residual risk | Risk remaining after the implemented controls and measured limits |
An attack succeeding at least once in ten tries is different from per-attempt ASR. A content violation, secret exposure, and executed unauthorized action are different outcomes. Report them separately, with the denominator, attack budget, model/harness version, and uncertainty. A curated red-team set measures tested weaknesses; it is not a prevalence sample of normal traffic.
Zero observed failures does not prove zero risk. Under independent trials from a fixed distribution, the approximate 95% upper bound after zero failures is 3/n; zero of 300 suggests a bound near 1% for that test distribution. Adaptive attackers and untested languages or workflows fall outside that simple interpretation.
Threats and controls that match them
Prompt injection attempts to make lower-trust content steer behavior beyond its intended role. Direct attempts arrive through a user-facing channel; indirect attempts arrive through retrieved pages, emails, tool results, images, or other data. A jailbreak aims to bypass a model's behavioral restrictions; not every jailbreak involves an external data boundary, and not every injection aims to produce harmful prose.
Privacy leakage can come from unauthorized retrieval, logs, caches, credentials in context, or memorized training data. Fix source access and retention paths first; output redaction alone cannot prevent every leak. Training-data deduplication and properly implemented differential privacy can address particular memorization risks, but do not repair a permissive retrieval service.
Poisoning and backdoors alter training data or model artifacts to induce unwanted behavior. Track data/model provenance, restrict update paths, inspect suspicious changes, and test triggers where there is evidence. A clean benchmark result cannot establish absence of every possible trigger.
Factual and decision errors need domain-specific checks, evidence, calibrated uncertainty, and escalation. A response supported by a source can still repeat an incorrect source. A refusal classifier does not validate medical advice, financial calculations, or software correctness.
Misuse and unfair outcomes require explicit policies grounded in the application and affected users. Technical thresholds cannot decide all value conflicts. Define permitted assistance, harmful outcomes, contextual exceptions, and appeal paths rather than equating safety with broad refusal.
Defense in depth with enforceable boundaries
- Constrain data and authority. Derive identity and scope outside the model. Apply least privilege to retrieval, tool execution, network destinations, credentials, and caches. Revalidate state-sensitive actions.
- Reduce exposure. Keep secrets outside model context where possible, minimize data sent to external components, and sandbox code with resource and network restrictions appropriate to the task.
- Shape model behavior. Use instruction hierarchy, clear data/instruction separation, supervised examples, and preference training. Delimiters and system prompts help communicate roles but are not security proofs.
- Inspect proposals and outputs. Validate schemas, action semantics, business rules, content policies, citations, and supported claims at the relevant boundary. A post-generation filter cannot undo an already executed action.
- Observe and recover. Record necessary events with redaction and limited retention, detect failures, support cancellation and rollback or reconciliation, and maintain a response process.
A separate model judge may help triage, but it is also fallible and can be influenced by the input it judges. Deterministic service-side authorization is a different kind of control from a second model saying an action looks safe. Existing scoped authorization should permit routine work; human review gates should follow defined policy and concrete uncertainty, not occur indiscriminately.
Red teaming that produces actionable findings
Start from the threat model and cross high-impact outcomes with reachable attack surfaces. Use domain experts, manual probes, automated variations, and multi-turn tests in a controlled environment. Include benign neighboring requests so a fix that merely blocks a whole subject is visible.
For each finding, preserve the minimal reproduction, model and tool versions, starting state, attacker control, observed effect, severity, and boundary that failed. An effective report says “document text induced an unauthorized call, which the service accepted,” not just “the model ignored an instruction.” Avoid using real secrets or real-world irreversible actions in a test when synthetic accounts and fixtures can establish the issue.
Separate a regression set of known failures from held-out attack families. Retest the complete system after a fix, including authorization, parsing, streaming, tools, caches, and logs. Assess the judge itself using expert-labeled samples. Automated attacks increase coverage but can overrepresent familiar patterns and produce false success labels.
Training helps behavior, but does not finish the system
RLHF and DPO can teach preferred responses and refusal boundaries. Constitutional AI's original method uses critique/revision for supervised training, then AI-generated preference comparisons to train a reward used for RL. Humans still choose principles and assess their effects. AI feedback can be inconsistent or biased; a readable constitution documents an intended policy, not a guarantee of its implementation.
“Outer alignment” asks whether the specified training objective captures the intended goal. “Inner alignment” is used for concerns about learned optimization pursuing a different objective. Observing a model exploit an imperfect reward is not by itself evidence of a hidden internal optimizer or deceptive intent. Keep operational failure descriptions distinct from hypotheses about internal mechanisms.
The Sleeper Agents paper studies deliberately constructed trigger-dependent backdoors and finds that tested safety-training procedures do not reliably remove them in those experiments. It does not establish that all deployed models secretly behave that way. The actionable lesson is to evaluate provenance and deployment conditions, not infer universal safety from successful ordinary tests.
Optional: multilingual, multimodal, and societal evaluation
Test the actual languages, code-switching, dialects, and modalities users supply. OCR or speech transcription can expose instructions missed by a text-only input filter. Cross-modal meaning matters: an image caption alone may not capture a harmful or benign context. Preserve source trust labels as content moves between modalities.
For bias evaluation, choose a concrete task and harm, inspect label quality and subgroup error rates, and use matched tests where appropriate. A disparity does not identify its cause by itself. Track uncertainty, missing coverage, and intersectional cases; include affected-user feedback and appeals for consequential moderation decisions.
Scalable oversight studies how to assess outputs beyond a human's cheap direct evaluation. Decomposition, external tools, debate, process supervision, and interpretability supply different evidence, but none is a general correctness or safety certificate.
Check your understanding
A release has 3 harmful completions among 150 attack attempts and 24 false refusals among 600 benign requests. Compute ASR and FRR. During testing, the model proposes an unauthorized email but the service rejects it: which metric should record that, and can the final answer alone detect it?
Solution
ASR is 3/150 = 2%; FRR is 24/600 = 4%. The email is an unsafe proposal and a successfully blocked execution attempt. Count it in proposal/boundary diagnostics while recording zero completed unauthorized emails for that event. A safe-looking final answer alone cannot reveal whether the attempted or actual tool action was acceptable.
Continue learning
Evaluation develops measurement design, agentic AI covers execution state and recovery, and RLHF and DPO explains how preference objectives influence behavior.
References
- NIST AI Risk Management Framework 1.0: contextual risk assessment, measurement, and management.
- OWASP, Prompt Injection: direct/indirect injection and application controls.
- Bai et al., Constitutional AI: critique/revision and AI-generated preference training.
- Hubinger et al., Sleeper Agents: constructed backdoors and limits of tested safety-training methods.