Anthropic RSP 3.4: AI risk thresholds and governance
Responsible Scaling Policy 3.4 changes R&D thresholds and Risk Report rules. Translate Anthropic's model into evidence-based enterprise AI governance.
- AUTHOR
- Karol Rapacz / Penetration Tester (OSCP, PNPT)
- PUBLISHED
- 23 June 2026
- READING TIME
- 10 min read
- TOPIC
- AI Governance
Anthropic Responsible Scaling Policy 3.4 took effect on 8 July 2026 and describes governance for risks that may emerge as model capabilities grow. The update changes the automated R&D threshold, Risk Report publication rules and external-review structure.
RSP is one supplier’s voluntary policy, not a standard or proof that a deployment is safe. It can nevertheless provide a useful threshold- and evidence-based governance pattern.
What version 3.4 changed
Anthropic lists five changes:
- revise the automated R&D threshold to track the threat model more closely;
- distribute complete Risk Reports internally to at least 200 employees instead of all regular-clearance staff;
- permit reports to analyse risk as of an explicit coverage date;
- require public reports to indicate where material was redacted;
- allow several external reviewers when every unredacted section is assessed by at least one reviewer.
These changes illustrate the tension between currency, confidentiality and independent oversight.
Capability thresholds, not one model score
A threshold should connect to a harm scenario. A model may be excellent at coding without crossing an automated-research or biological-risk threshold. One “high risk” label hides these differences.
Enterprises should define domain-specific thresholds: production autonomy, data access, financial operations, cyber capability and influence over human decisions.
Risk Reports as decision artefacts
A report should identify system version, coverage date, scenarios, tests, results, uncertainty, safeguards and residual-risk owner. It must lead to a decision: deploy, constrain, retest or stop.
Marking redactions tells a reader that evidence exists but is withheld. It does not allow verification, so critical sections need trusted review under NDA or another controlled route.
Multiple reviewers
Splitting a report among experts is sensible when a chemist should not assess exploit development and a red teamer should not evaluate biological methodology. Complete coverage and an integrator for conflicting conclusions are essential.
Maintain a matrix of section, required expertise, reviewer, conflict of interest, date and outcome. One unreviewed critical appendix cannot disappear inside an overall “approved” status.
Off-cycle review after change
Versions 3.2 and 3.3 strengthened external review and model-specific risk updates. The practical lesson is that a quarterly calendar is insufficient. A new model, tool, RAG source, autonomy level or retention policy can require immediate assessment.
Connect change records with the AI model registry. Each material change should trigger the relevant evaluation suite.
Enterprise implementation
- Define harm scenarios and measurable thresholds.
- Assign business and technical owners.
- Establish tests below and above thresholds.
- Produce a versioned Risk Report with a coverage date.
- Require review by appropriate experts.
- Record redactions and the location of complete evidence.
- Define controls required after a threshold is crossed.
- Monitor changes and trigger off-cycle assessment.
Limitations
A policy does not prevent incidents by itself. Thresholds can be miscalibrated, evaluations can miss production behaviour and reviewers may lack context. Logs, incident response, red teaming and observation of real effects remain necessary.
RSP 3.4 is valuable as an example of a transparent update mechanism. The key enterprise lesson is that governance must react to capabilities and system changes, not merely model names or annual questionnaires.
Reading a governance commitment
An RSP is one company’s policy, not an independent standard or guarantee that a model is safe. Analyse the exact version, threshold definitions, required evaluations, decision authority and exception process. Separate a public commitment from a technical control you can verify in your deployment.
For a customer, practical questions are which model version runs, which safeguards are active, where evaluation results exist and what happens when risk classification changes. Retain the document version used in supplier assessment because policies evolve. Your own assessment still covers data, tools, users and application failure impact.
Sources: Anthropic RSP, RSP Version 3.4, Frontier Safety Roadmap.


