**Pre-training evaluation.** Before a large training run, Anthropic estimates the model's expected capabilities based on scaling laws + smaller-scale ablations. If the forecast indicates crossing into a higher ASL without mitigations in place, the RSP commits to pausing the run.
**Mid-training evaluation.** During large training runs, intermediate checkpoints are evaluated. If capabilities are tracking faster than forecast, training can be paused early.
**Pre-deployment evaluation.** Final-checkpoint evaluations including red-team, capability evals (bio, cyber, autonomy, persuasion), and external evaluator runs (UK AISI, US AISI, METR, Apollo). Findings inform the Capability + Safeguards Reports.
**Threshold-crossing decision.** If evals indicate crossing into a higher ASL, the Responsible Scaling Officer prepares the case. The CEO signs off on deployment decisions involving newly-crossed thresholds. The Board and the Long-Term Benefit Trust have oversight authority.
**External evaluator role.** UK AISI, US AISI, METR, Apollo Research, and other external evaluators provide independent capability assessments. Their findings inform Anthropic's internal threshold decisions; selective findings appear in the published Capability + Safeguards Reports.
**Pause provision.** If a threshold crossing is identified without corresponding mitigations being ready, the RSP commits to pausing further training until mitigations are in place. As of June 2026, no documented public pause has been cited — which could mean (a) the framework is upstream-shaping (development pace matches mitigation pace), or (b) the framework hasn't yet been stress-tested. Independent observers (UK AISI, US AISI, METR, Apollo, academic researchers) will continue to probe this.