Skip to contentNew: Does ChatGPT recommend your brand? Free 60-second AI visibility check →
By The AI Prompts Hub Team · Digital Empire

Anthropic RSP ASL Levels Explained (2026): ASL-1 to ASL-5

By DDH Research Team at Digital Dashboard HubUpdated

Stop writing AI prompts from scratch.

Tell us your business + your task + your model. We write the prompt — perfectly tuned for ChatGPT, Claude, Grok, Gemini, Midjourney, or any model. Plus 500+ pre-built prompts in your library.

14 days, no card. Cancel in 2 clicks.

Anthropic's Responsible Scaling Policy (https://www.anthropic.com/rsp) is the company's public commitment that frontier-model training and deployment will gate on capability evaluation. The unit of risk is the **AI Safety Level (ASL)** — a tiered scheme modeled on the BSL (biosafety level) tiers used in biological research. Each ASL carries deployment + security commitments that must be in place before a model crossing into that level can be trained further or deployed.

Five levels are defined. **ASL-1**: systems with no meaningful risk (smaller-than-frontier models, narrow systems). **ASL-2**: systems with 'early signs of dangerous capabilities' (current frontier chat models). **ASL-3**: substantially elevated misuse risk OR low-level autonomous capabilities (current Claude Opus 4 / 4.7 on specific axes). **ASL-4**: substantially more capable systems requiring mitigations Anthropic has been building toward. **ASL-5**: substantially super-human systems with mitigations Anthropic explicitly states it has not yet developed.

The RSP commits Anthropic to: (a) evaluating models pre-deployment + during training for ASL-relevant capabilities, (b) implementing the corresponding ASL's mitigations before crossing the threshold, (c) pausing further training if pre-training evals forecast crossing into a higher ASL without the corresponding mitigations in place, (d) publishing Capability + Safeguards Reports for ASL-3+ models documenting what was evaluated + what mitigations are deployed.

Companion guides: Anthropic RSP vs OpenAI Preparedness, OpenAI Preparedness Framework Thresholds, Superalignment vs RSP vs Frontier Safety.

Digital Dashboard Hub

Rate limits hurt because prompts are loose — ITPM blows up before RPM does. DDH's AI Prompt Builder writes cache-anchored prompts so 80%+ of your input tokens are billed at 10% of the rate, and you hit limits later (or never).

Start free 14-day trial — AICHAT30 = 30% off Pro for 3 months.

ASL ladder summary — June 2026

Feature
Description
Current models at this level
Deployment standard
Security standard
ASL-1No meaningful catastrophic-risk capabilityNarrow systems, classifiers, very small modelsStandard practicesStandard practices
ASL-2Early signs of dangerous capabilitiesMost frontier chat models including Claude Haiku 4.5, Sonnet 4.6 baseline behaviorLimited release with refusal training + acceptable-use policiesIndustry-standard cloud + access controls
ASL-3Substantially elevated misuse risk OR low-level autonomous capabilityClaude Opus 4 + Opus 4.7 on specific axes per published Capability ReportsEnhanced misuse protections + sandboxed deployment + third-party Capability Report attestationHardened security against opportunistic attackers; insider-threat controls
ASL-4Significantly more capable; meaningfully elevated autonomy or substantial misuse potentialNo deployed models yet at June 2026Mitigations Anthropic has been building toward; deployment requires additional safeguardsSecurity against state-level adversaries; advanced insider-threat detection
ASL-5Substantially super-human capabilitiesNo deployed models yet at June 2026Mitigations Anthropic explicitly states it has not yet developedMitigations Anthropic explicitly states it has not yet developed

Source: Anthropic Responsible Scaling Policy at https://www.anthropic.com/rsp (v2.0 October 2024 + subsequent updates through 2025-2026). Capability + Safeguards Reports for Claude Opus 4 and Opus 4.7 published at anthropic.com/news. ASL definitions are deliberately broad in advance of capability rather than precisely calibrated to specific eval scores; Anthropic explicitly commits to crossing thresholds only after the corresponding mitigations are demonstrated.

ASL-1: no meaningful catastrophic-risk capability

**Definition.** Systems that pose no meaningful catastrophic-risk capability. Narrow systems, classifiers, simple language models below frontier capability. Most production ML systems pre-2022 sit at ASL-1.

**Deployment standard.** Standard industry practices. No RSP-specific deployment requirements.

**Security standard.** Standard industry practices (standard cloud security, standard access controls).

**What's at ASL-1 today.** Image classifiers, narrow speech-recognition, recommendation systems, smaller-than-frontier language models, fine-tuned models for narrow business tasks. Most enterprise ML deployments.

**Significance.** ASL-1 is the baseline. Most of the ML world operates here. RSP-specific obligations only apply when a system rises to ASL-2 or above.


ASL-2: early signs of dangerous capabilities

**Definition.** Frontier-quality language models that show early signs of dangerous capabilities — uplift on misuse vectors compared to lower-capability baselines, but below the threshold where catastrophic risk is materially elevated.

**What's at ASL-2 today.** Most current frontier chat models including Claude Haiku 4.5, Claude Sonnet 4.6 on most axes, similar-capability releases from OpenAI / Google / others (each lab uses its own framework, but the capability range is roughly comparable). Models in this range pass standard evaluation for harmful-content elicitation, prompt injection resistance, and basic misuse-resistance, but may show concerning capabilities on adversarial probes.

**Deployment standard at ASL-2.** Limited release with refusal training (the model is trained to refuse high-risk requests), acceptable-use policies (developers + end-users agree to policy terms), and operational safeguards (rate limits, content filtering, monitoring).

**Security standard at ASL-2.** Industry-standard cloud security (encryption at rest + in transit, role-based access control), perimeter security, employee access controls + auditing. Comparable to mature SOC 2 / ISO 27001 security posture.

**Significance.** ASL-2 is where most current frontier chat models sit. The deployment + security commitments are non-trivial but match what mature SaaS companies routinely operate. ASL-2 represents 'normal frontier-model operations' in 2026.


ASL-3: meaningfully elevated risk + hardened mitigations

**Definition.** Systems where capability evaluations indicate (a) substantially elevated risk of catastrophic misuse (uplift on bio/chem/cyber/etc. attack capability compared to current baselines), OR (b) low-level autonomous capability (the model can plausibly take coherent multi-step actions toward a goal across realistic environments).

**Two axes — misuse + autonomy.** A model can cross into ASL-3 on the misuse axis without crossing on the autonomy axis, or vice versa. Anthropic's published Capability Reports for Claude Opus 4 and 4.7 indicate the model is ASL-3 on specific misuse-related axes and approaches but has not crossed the autonomy axis at ASL-3.

**Deployment standard at ASL-3.** Enhanced misuse protections: stronger refusal training, additional safety-classifier layers, restricted access patterns (no anonymous high-volume queries, organizational accountability), sandboxed deployment for higher-risk use cases, third-party Capability Report attestation (published Capability + Safeguards Reports prepared with UK AISI / US AISI / METR evaluator input).

**Security standard at ASL-3.** Hardened security against opportunistic attackers: enhanced perimeter security, multi-party access controls for model weights (no single-person access to the full weights), monitoring + audit logging of weights access, insider-threat controls. Designed to defeat opportunistic + low-resource attackers (script kiddies, opportunistic insiders, non-state criminal actors).

**Published Capability + Safeguards Reports.** Anthropic publishes per-model artifacts for ASL-3 releases. The Capability Report documents what evaluations were run + headline findings. The Safeguards Report documents what mitigations are in place + why they are considered adequate for the model's capability profile. Both shipped for Claude Opus 4 and Opus 4.7 (linked from anthropic.com/news).

**Significance.** ASL-3 is where the RSP becomes operationally visible. The published Capability + Safeguards Reports are the artifact procurement and compliance teams should pull for diligence on ASL-3 models. The deployment + security commitments require substantial investment to implement, which is why ASL-3 release is treated as a distinct milestone.


ASL-4: significantly more capable + state-level security

**Definition.** Substantially more capable models — meaningfully elevated autonomy (persistent agentic behavior across realistic environments), substantial misuse potential (uplift to expert teams attempting catastrophic attacks), or AI R&D capability (the model meaningfully accelerates AI research, raising recursive-self-improvement concerns).

**Current status at June 2026.** No deployed models at ASL-4. Anthropic's RSP commits the company to having ASL-4 mitigations in place before training or deploying ASL-4 models — implementation work is ongoing.

**Deployment standard at ASL-4 (planned).** Substantially stronger deployment-misuse mitigations. Vetted-user-only access for higher-capability axes. Real-time monitoring of agent actions. Limited tool-use surfaces. Significantly tighter capability disclosure (the model card / Capability Report may have more redactions for sensitive details).

**Security standard at ASL-4 (planned).** Security against state-level adversaries. Compartmentalized access (multi-party + multi-region access controls on weights). Advanced insider-threat detection. Air-gapped or near-air-gapped training environments. Match the security posture of organizations handling state-secret-level information.

**Implementation challenge.** The ASL-4 commitments require new operational infrastructure that did not exist for ASL-3. Building this is the substrate of Anthropic's safety + security team work through 2024-2026. Crossing into ASL-4 without these mitigations is the line Anthropic has publicly committed not to cross.

**Significance.** ASL-4 is the level where commercial deployment shape changes materially. Procurement, compliance, and trust questions become harder. The published Capability + Safeguards Reports will likely have more redactions; third-party evaluator engagement will be substantially deeper. Industry will likely face EU AI Act systemic-risk classification implications + sector-regulator scrutiny that doesn't yet exist for ASL-3 deployments.


ASL-5: super-human capability + un-developed mitigations

**Definition.** Substantially super-human capabilities — models that meaningfully exceed human ability across cognitive domains. Anthropic's RSP explicitly states that ASL-5 deployment requires mitigations the company has not yet developed.

**Current status at June 2026.** No models at ASL-5. ASL-5 is specified in the RSP as a placeholder — Anthropic commits in advance to not deploying ASL-5 systems without the corresponding mitigations, even though the company does not yet know exactly what those mitigations look like.

**Why include ASL-5 in the policy.** Two reasons. (1) **Forcing function**: by specifying ASL-5 as a level with un-developed mitigations, Anthropic publicly commits to not racing past the line. (2) **Research direction**: explicitly naming ASL-5 frames the open research problems (scalable oversight, weak-to-strong generalization, deceptive alignment detection) as requirements not options.

**What ASL-5 mitigations might look like.** Open research questions: alignment techniques that scale to super-human capability (where evaluators cannot directly verify the model's intent), interpretability tools that can audit models more capable than the auditor, governance structures that can credibly halt deployment under commercial pressure. Anthropic, OpenAI, DeepMind, academic groups, and external safety organizations all actively research these questions.

**The honest framing.** ASL-5 is the limit case. Whether models capable enough to require ASL-5 mitigations actually exist by 2027, 2030, 2040, or never is a matter of forecasting that reasonable people disagree on. The RSP's ASL-5 placeholder is a public commitment that — whenever such models emerge — Anthropic will not deploy them without first developing the corresponding mitigations. That commitment is the substantive content of ASL-5, more than the technical specification.


How crossing thresholds actually works

**Pre-training evaluation.** Before a large training run, Anthropic estimates the model's expected capabilities based on scaling laws + smaller-scale ablations. If the forecast indicates crossing into a higher ASL without mitigations in place, the RSP commits to pausing the run.

**Mid-training evaluation.** During large training runs, intermediate checkpoints are evaluated. If capabilities are tracking faster than forecast, training can be paused early.

**Pre-deployment evaluation.** Final-checkpoint evaluations including red-team, capability evals (bio, cyber, autonomy, persuasion), and external evaluator runs (UK AISI, US AISI, METR, Apollo). Findings inform the Capability + Safeguards Reports.

**Threshold-crossing decision.** If evals indicate crossing into a higher ASL, the Responsible Scaling Officer prepares the case. The CEO signs off on deployment decisions involving newly-crossed thresholds. The Board and the Long-Term Benefit Trust have oversight authority.

**External evaluator role.** UK AISI, US AISI, METR, Apollo Research, and other external evaluators provide independent capability assessments. Their findings inform Anthropic's internal threshold decisions; selective findings appear in the published Capability + Safeguards Reports.

**Pause provision.** If a threshold crossing is identified without corresponding mitigations being ready, the RSP commits to pausing further training until mitigations are in place. As of June 2026, no documented public pause has been cited — which could mean (a) the framework is upstream-shaping (development pace matches mitigation pace), or (b) the framework hasn't yet been stress-tested. Independent observers (UK AISI, US AISI, METR, Apollo, academic researchers) will continue to probe this.


Implications for your team

**Procurement diligence.** If you're licensing Claude Opus 4 / 4.7 for a regulated industry, pull the published Capability + Safeguards Reports. They document what was evaluated + what mitigations apply. The reports are part of your AI provider diligence file (alongside SOC 2, ISO 27001, ISO/IEC 42001 attestations).

**ASL-3 = enterprise readiness.** ASL-3 models are designed for enterprise deployment with mature procurement processes. The published reports + the deployment + security commitments are the substrate of the enterprise relationship.

**Vendor portability.** All major frontier labs operate under voluntary safety frameworks that reserve the right to restrict access if thresholds are crossed. Design for portability via abstraction. If you use Claude Opus 4.7 (ASL-3) heavily, plan for the possibility of access restrictions if Anthropic identifies a finding requiring deployment changes.

**Comparing across labs.** Anthropic uses ASL; OpenAI uses Tracked Categories × thresholds; DeepMind uses CCLs per domain. Different organizing concepts; broadly comparable practical decision tree. See RSP vs Preparedness and Superalignment vs RSP vs Frontier Safety for the comparative analysis.

**For your own AI products.** Adapting the ASL framing to your application (a 'safety level' scheme for your specific use cases with explicit deployment + monitoring commitments per level) is a useful internal governance pattern even if your product is far below frontier capability. The methodology is portable; the specific thresholds are application-specific.

Working with ASL classifications

  1. 1

    Pull the Capability + Safeguards Reports for your candidate model

    For Claude Opus 4 / 4.7 (ASL-3 models): linked from anthropic.com/news. Each report documents evaluation methodology + headline findings + mitigations. Read in full before procurement.

  2. 2

    Read the full RSP text

    anthropic.com/rsp. ~20-40 pages. Revision history, definitions, commitments, threshold descriptions. The source of truth — anything written about ASL is a digest of this document.

  3. 3

    Cross-reference with third-party evaluator reports

    UK AISI (aisi.gov.uk/work/our-publications), US AISI (aisi.nist.gov/publications), METR (metr.org/blog), Apollo Research (apolloresearch.ai/research). Their methodology + findings provide independent perspective on the RSP's effectiveness.

  4. 4

    Plan for vendor portability in your architecture

    RSP gives Anthropic the right (and commitment) to restrict access if thresholds are crossed. Design abstraction layer over the Anthropic + OpenAI + Google SDKs. Config-driven model identifiers. Prompt formats portable across shapes.

  5. 5

    Adapt the methodology for your own application

    Even if your application is far below frontier capability, a 'safety level' scheme with explicit per-level deployment + monitoring commitments is useful internal governance. Pair with Build LLM Red-Team Suite 2026 and Implement Constitutional AI Guardrails.

    → Open the Implement Constitutional AI Guardrails

Use the data programmatically

Every page on this site is also exposed as a free, CORS-open JSON endpoint. No auth, no rate limit (fair-use, please cache). License is CC-BY-4.0 — link back to attribution.canonicalUrl in the response.

Endpoint: https://aipromptshub.co/api/limits/anthropic-rsp-asl-levels-explained
curl
curl -s 'https://aipromptshub.co/api/limits/anthropic-rsp-asl-levels-explained' | jq .
Python
import requests

r = requests.get("https://aipromptshub.co/api/limits/anthropic-rsp-asl-levels-explained", timeout=10)
r.raise_for_status()
data = r.json()
print(data["title"])
for source in data.get("sources", []):
    print("source:", source)
JavaScript / Node
// Node 20+ / modern browser
const res = await fetch("https://aipromptshub.co/api/limits/anthropic-rsp-asl-levels-explained");
if (!res.ok) throw new Error("HTTP " + res.status);
const anthropic_rsp_asl_levels_explained = await res.json();
console.log(anthropic_rsp_asl_levels_explained.title);
for (const source of anthropic_rsp_asl_levels_explained.sources ?? []) {
  console.log("source:", source);
}

Spec: /api/openapi.yaml · Docs: /api/docs

Frequently Asked Questions

What is ASL-2?

ASL-2 is the AI Safety Level for systems with 'early signs of dangerous capabilities' but below the threshold of meaningfully elevated catastrophic-risk capability. Most current frontier chat models sit at ASL-2 — including Claude Haiku 4.5, Claude Sonnet 4.6 on most axes, and similar-capability releases from other labs. Deployment standard: limited release with refusal training + acceptable-use policies. Security standard: industry-standard cloud security + access controls.

What is ASL-3?

ASL-3 is the AI Safety Level for systems with substantially elevated risk of catastrophic misuse OR low-level autonomous capability. Current Claude Opus 4 and Opus 4.7 are at ASL-3 on specific axes per published Capability Reports. Deployment standard: enhanced misuse protections + sandboxed deployment + third-party Capability Report attestation. Security standard: hardened security against opportunistic attackers + insider-threat controls. Capability + Safeguards Reports are published for ASL-3 releases.

Which Claude models are at ASL-3?

Per published Capability Reports linked from anthropic.com/news: Claude Opus 4 (released 2025) and Claude Opus 4.7 (released 2026) are at ASL-3 on specific capability axes — particularly bio/chem-uplift evaluations on certain probes, and on cybersecurity capability evaluations. The reports are explicit about which axes are at ASL-3 vs ASL-2; not every axis crosses the threshold. Claude Sonnet 4.6 and Claude Haiku 4.5 are at ASL-2 across reported axes.

What's the difference between ASL-3 deployment and security standards?

Deployment standards govern how the model is rolled out to users: refusal training depth, access patterns, sandbox requirements, user/developer agreements, third-party attestation. Security standards govern how the model weights themselves are protected: encryption, multi-party access, insider-threat detection, perimeter security. ASL-3 raises both standards above ASL-2. ASL-4 raises both further, with security designed to defeat state-level adversaries.

Has Anthropic ever paused training because of an ASL forecast?

Per public communications as of June 2026, no specific instance of a documented pause has been cited. This could mean (a) the framework is upstream-shaping — Anthropic's research + development pace matches mitigation pace, so no pause has been needed, or (b) the framework hasn't yet been stress-tested by a borderline decision. UK AISI, US AISI, METR, Apollo Research, and academic safety researchers continue to probe this question.

How is ASL different from OpenAI's Preparedness thresholds?

Structurally: Anthropic uses a single overall ASL ladder; OpenAI uses Tracked Categories (Bio/Chem, Cyber, AI Self-Improvement, etc.) × Low/Medium/High/Critical thresholds per category. Practically: similar decision tree in the same risk scenario — document the capability, add mitigations before deployment, consult external evaluators, publish a per-model artifact. The structural difference matters most on borderline cases + on the strength of the governance backstop (Anthropic's Long-Term Benefit Trust is the most distinct structural feature). See Anthropic RSP vs OpenAI Preparedness Framework.

What are Capability + Safeguards Reports?

Per-model artifacts that Anthropic publishes for ASL-3+ models. The Capability Report documents what evaluations were run + headline findings. The Safeguards Report documents what mitigations are in place + why they are considered adequate for the model's capability profile. Both shipped for Claude Opus 4 and Opus 4.7. Includes summaries of external evaluator findings (UK AISI, US AISI, METR, Apollo) where applicable. Not published: raw eval scores, full red-team transcripts, specific failure-mode prompts.

Does the RSP apply to fine-tuned models?

Yes. The RSP applies to Anthropic's deployment of any Claude model variant; fine-tuned customer models built on Claude inherit the underlying model's ASL classification. Anthropic's safety screening + content filters apply to fine-tuned models — customer fine-tuning can refine helpful behavior but cannot disable core refusal training or safety filters. For ASL-3 base models (Opus 4 / 4.7), fine-tuned variants are also subject to ASL-3 deployment standards.

ASL is the policy. Your prompt design is where Claude's behavior ships.

Once you've procured an ASL-3 Claude model, the prompts you write determine whether the safety surface helps or fights your application. Our AI Prompt Generator writes Claude-tuned prompts that work with Constitutional AI behavior — cache-anchored, XML-structured, instruction-hierarchy-aware. 14-day free trial, no card.

Browse all prompt tools →