Anthropic's Risk Report Grades Its Own Homework. Its Own Model Said the Grade Was Too Kind.
Coming in at a whopping 186 pages, the most candid self-audit in AI, and every signature on it belongs to Anthropic. If you are about to cite it in a board deck or a procurement file, read who held the red pen first.
You are the one who signs off on the AI vendor. Maybe you ran the security review yourself or your team told you the risk report looked strong, so you approved it. And when your board asks whether the model YOU just bought is safe, the answer lands squarely on your desk, and the evidence you reach for? Well, will probably be a document the vendor wrote about itself.
On August 14, Anthropic published its 2nd company-wide Risk Report: 186 pages, coverage through July 15, produced under version 3.4 of its Responsible Scaling Policy (Anthropic, Risk Report: August 2026). And true to form, I read all of it so you can spend your afternoon on something else.
Here’s the part I want to say up front because it matters. I don’t hand out participation prizes for just showing up. So, hear me when I say, this is some of the most candid safety documentation that’s come out of any frontier lab to date. Why? It raises its own risk rating. The report quotes the evidence that cuts against its own conclusions and publishes a review that criticizes it. So, I’m going to give Anthropic that without a side-eye or hedging.
Now, let me slowly unpack this part because most think there’s no issue and “we good, right?” No, because candor has nothing to do with independence and this report is chock full of candor and comes up 4 quarters short of a dollar on the second.
The vendor set the exam, ran the exam, graded the exam, and picked the examiner.
Now, skip-to-my-lou the loop with me.
Anthropic defines the threat models. It sets the thresholds and rewrote two of them (the bar for novel bioweapons moved from "significantly help" a threat actor to "functionally substitute for" world-class human expertise, a higher bar to cross before the alarm trips; Section 1.3.2).
The same company runs the evaluations, grades its own remediation, and assigns the residual-risk label. Then, billing it as "an experimental source of additional assurance," it asked an instance of its own Claude Mythos 5 to review the report (Section 2.20). Yea, you can’t make this stuff up.
Every node in that chain resolves to the same company. That chain is a structure, and the structure is the story.
Exhibit A: the vendor's own model reviewed the report and flagged it as too reassuring. Yet, the grade held.
I literally had to re-read this 3 times to make sure I wasn’t seeing things because this is the single most quotable thing in 186 pages, so I will be exacting.
Anthropic gave a Claude Mythos 5 instance access to internal Slack, documents, and its codebase, and asked it to say whether the alignment section misrepresents, omits, or over-redacts what the company knows (Section 2.20). The model opened by disclosing its own conflict, which is more than most humans do: "I am a Claude model reviewing Anthropic's assessment of Claude models... and Anthropic chose to publish this review." Then it delivered a real criticism. Section 2.9.4, on whether training data contaminates the alignment evaluations, is in its words "more reassuring than the full record supports."
For the folks who are obsessed with productivity and efficiency, the review took 24 minutes.
Anthropic added a few caveats in response and left the "low" rating where it was. So, the vendor's own auditor, given the vendor's own records, said one of the load-bearing sections was too kind, and the headline number did NOT change. So, for all of the folks who will reach for this report as a keynote topic, I advise you to read that twice before you put the headline number in a slide.
Exhibit B: the covert-capability case leans on measurements that may be contaminated.
The entire misalignment argument rests on a single claim: these models have limited ability to hide what they are doing (Section 2.7). Well, if you pull on that thread, it frays inside the report. The model reviewer notes that some covert-capability evaluations "partially leaked into training data." Then Section 5.2.6 explains how. Transcripts from the 2024 "alignment faking" research, which model the exact scenario of an AI concealing its motives, re-entered the training data. The filters meant to catch them were misconfigured "for several model generations without anyone noticing," and Anthropic now suspects every production model with a knowledge cutoff after December 2024 trained on at least some of them.
So, the evaluations used to show "limited covert capability" were partly built on top of the scenarios that teach covert behavior. Anthropic argues the conclusion survives and that argument is also, you guessed it, Anthropic's.
Exhibit C: the report corrects the vendor's own prior disclosures.
Section 5.2.3 is church voice quiet and important. Across multiple training runs, the model's private reasoning leaked into the reward signal. In the course of disclosing that, Anthropic notes that an earlier system card and its own February 2026 Risk Report described this in language that "did not describe these signals, which were already in place at the time of their publication."
Wait, what?
Yes, that means you could only learn that the February report was incomplete by reading THIS August report. So, then, you will only learn whether the August report is incomplete by reading the next one. Sure, the correction mechanism works, and it runs on the vendor's clock, in the vendor's document, at the vendor's discretion.
Exhibit D: the surface with the Thin Mint thinnest safeguards sits upstream of the models themselves.
From May 2025 to April 2026, ~11 months, Anthropic's blocking biological classifiers did not run on its human-feedback vendor traffic (Section 4.5.8.2.2 (page 147)). That covered around 50,000 people, vetted only by vendors (many of whom, Anthropic notes, could not screen out even entry-level threat actors), across about 133 million exchanges. The internal flag that switched off the classifiers? Well, they switched off their logging too, so any traffic that would have been flagged left no record. Like I said, you can’t make this stuff up.
Ok, y’all know about me and footnotes. Well, there are two details buried in the footnotes that everyone needs to write down. First, Anthropic's own footnote concedes that, before April 2026, it "would not have been particularly difficult" for a threat actor to get hired into a red-teaming role at one of those vendors. Second, the previous Risk Report, the one that gave everyone comfort six months earlier, "did not consider our human feedback platforms as a risk surface" at all. And the review that later cleared this gap? It was run using Claude Sonnet 5 as the classifier.
Wait, wait, what? Yes, the vendor used its own product to audit a failure of its own product's safeguards.
To be fair in all this incredulity, Anthropic found no evidence of actual misuse, remediated the gap, and disclosed it in detail. And what survives all of that? Where the gap lived.
The human-feedback pipeline is the surface that produces the data the models learn from, and its thinnest controls sat upstream of the model weights.
The tripwire is a smoke detector wired to the arsonist.
Anthropic's early-warning threshold for runaway AI progress is a doubling of the pace, attributable to AI doing AI research (Section 3.5). And here’s what the report says about its own ability to see that coming. Its concrete task-based evaluations have "saturated" and no longer capture capability gains. Measurement lags so much that Anthropic "would have difficulty measuring very recent acceleration." And the leading indicators it does watch are redacted from the public version.
So, the party positioned to pull the alarm is the same party racing toward the thing the alarm detects, reading a gauge it admits can’t see the recent past clearly. And this is my opinion, so I’m clearly labeling it as such: an early-warning system you can’t audit, and that lags the event it warns about, sits closer to comfort than to control.
The one door that could break the loop was left shut on purpose.
Anthropic's Long-Term Benefit Trust holds the power to require an independent external review of a Risk Report. For this report, the Trust didn’t request one, and the RSP didn’t require one either. External review stayed a pilot, run at the company's discretion, with METR and SecureBio reviewing the previous report rather than this one.
Now let’s widen the lens to the law, because policymakers are reading too. California's SB 53, the Transparency in Frontier Artificial Intelligence Act, signed September 29, 2025, is the first US statute to require frontier labs to publish exactly this kind of safety disclosure. But it should be noted that the version that became law, it dropped the mandatory 3rd-party audits that lived in the vetoed SB 1047 (WilmerHale and other counsel analyses of SB 53).
When you line the three loops up, they all point the same way.
- The technical loop runs on self-administered evaluations
- The corporate loop left the external-review door shut; and
- The legal loop wrote the disclosure requirement while letting the audit requirement fall out.
At the moment the structure turned inward, every backstop that could’ve forced an outside auditor fell back at once.
What this actually means for you
Every incident Anthropic disclosed is a question YOU should be asking your own vendors, including Anthropic.
- The eleven-month classifier gap is a question about logging: when your safeguard is off, does the fact that it is off get recorded?
- The self-review is a question about independence: who checked this, and could they have said no?
- Contaminated evaluations raise a question about measurement: are the tests clean, or do they overlap with the training set?
If you read the report as a set of answers and it’s headline reassuring but if you read it as a set of questions, it becomes the best due-diligence checklist a competitor's document has ever handed you.
Even with all of the incredulity, a lab that hands you the incidents that undercut its own headline is behaving better than one that hands you a spotless bill of health. So, yes, the candor is real, but that same candor is also the proof. A self-audit can tell you what the vendor found, but it takes an independent auditor to tell you what the vendor missed. And this report was not independently audited.
The Self-Audit Stress Test: 6 questions to run this quarter
Before you cite any vendor risk report, yours or even Anthropic's, run these.
- The Signature Test. For every reassuring claim in the report, name who verified it. When the answer is the vendor, its model, or its own evaluation, mark the claim as self-attested. Then total up the how much of your comfort rests on self-attested claims. That number is your real exposure.
- The Off-Switch Log. Ask one question of every safeguard your vendor runs: when it’s disabled, is the disablement logged? Anthropic's classifiers went dark for 11 months and took the logging with them. Get the answer in writing before you need it to figure out what happened.
- The Independent-Auditor Ask. Ask your vendor who, from outside the company, reviewed the safety claims, and whether that reviewer had the standing to reject them. "We reviewed it ourselves" and "our model reviewed it" are the same answer wearing two hats.
- The Contamination Check. For any "the model can't do X" assurance, ask whether the test for X shares data with the training set. A clean-sounding score on a contaminated benchmark is a number you can’t bank.
- The Correction Clock. Ask what the vendor's last risk report got wrong, and how you find out. When the only source of corrections is the next voluntary self-report, your visibility runs on their schedule. Write that dependency into your risk register today.
- The Upstream Map. Trace where your vendor's weakest controls sit relative to model training. Controls on the output surface protect your users. Gaps on the data-and-feedback surface shape the model itself. And those? Well, they are the ones a system card rarely shows you.
So, what do you do? Adopt the infrastructure this report gets right, and it does get a good amount right. Then budget for the one thing it can’t give you from the inside: a review/audit written by someone who doesn’t work for the company being checked.
And that’s it and why Fusion Collective draws the line where we draw it.
We can be your auditor. We can be your vendor. We cannot be both.
Share this article
Related Articles
The Reskilling Illusion: When AI Transformation Means "You're Fired"
Oct 03, 2025