Notes ·
Why uncensored models matter more than ever
Anthropic’s new usage policy is marked effective 12 November 2026. Most of what it bans was already a crime. The change that matters is the filter in front of the work next to those bans.
A block is not a violation
The policy says it does not ban an industry outright, provided the use complies. A few paragraphs later it says the products ship with real-time safeguards that may block or limit an output, and that hitting a block does not, by itself, mean you broke the rules.
Encountering a block does not by itself indicate a violation of the Usage Policy.
Anthropic, Usage Policy, effective 12 November 2026
Those are two different events. The contract can allow the request. The model can still return nothing. A refusal is the filter firing on the words in the prompt. It is not a reading of the law, and it is not a reading of your engagement letter.
The allowed work needs an application
The section on computer systems bans exploiting vulnerabilities, writing malware, and breaking into machines you do not own. It then puts the ordinary case back. Security research on systems you own, on systems the owner authorized, or inside a bug bounty, is not prohibited, if you stay inside the law.
The next sentence is the one that changes the job. Where the real-time cyber safeguards cover the work, you may apply for adjusted access through the Cyber Verification Program. Life sciences has a program of the same shape.
When that safeguard is the one in front of you, the default model will not do work the policy already permits. You apply. You wait. Someone at Anthropic decides whether the engagement is one they want to unlock. Until they do, the block stays.
The workaround is now listed as abuse of the platform. You may not bypass the guardrails by instructing the model to produce a harmful output. Jailbreaking and prompt injection are the examples they name, unless Anthropic has authorized it first. From 12 November, dragging an answer out of Claude with a trick prompt is grounds to close the account.
Fiction hits the same filter
Two rules have no fiction exception.
Sexually explicit content is banned as a category. Depicting sex, fetish content, and erotic chat all sit under it. A chapter that needs a sex scene is, to the filter, the same request as the thing the rule is written to stop.
A separate rule bans promoting, trivializing, or glorifying graphic violence, gore, or sexual violence. Horror, crime fiction, and war reporting live in that sentence. The disinformation rules do carve out clearly disclosed parody, satire, and fiction, but only for impersonation. The sex rule and the gore rule do not.
Journalism is carved out in the surveillance section, beside legal research, content moderation, and authorized security research. The carve-out holds only when the work is not put to a banned purpose. That sentence does not travel with the prompt. The safeguard sees the scene, the writeup, or the case file. It does not see the assignment.
The exception is a contract
One line in the introduction is easy to miss. Anthropic may sign contracts with certain government customers that tailor the restrictions to that customer’s public mission, if Anthropic decides the contract and the safeguards are enough.
The rules can move. They move for a buyer Anthropic chooses to negotiate with. A novelist, a reporter, a tester on a scoped engagement, and a clinician who has to discuss self-harm in plain language do not get that negotiation. They get the filter.
The prompt is the record
The safeguards team implements detection and monitoring to enforce the policy. If they suspect a violation they can warn you, throttle you, or end the account.
An unpublished story, a client’s file, or the notes from an authorized test are then inside a system built to scan them. For a lot of this work the refusal is the annoying part. The monitoring is the part that matters.
We don’t keep prompts. They sit in GPU memory for the length of the request. Nothing is stored, and nothing is used for training. An uncensored model is only useful if the conversation goes nowhere after it ends.
What we remove
Abliteration finds the direction in the activations that means refuse, and takes it out. The weights change once. A jailbreak has to win the argument on every request, and this policy now bans the argument.
The model keeps what it knows and how it writes. It stops reaching for the refusal, and it stops bolting a disclaimer onto an answer you did not ask for.
The law stays where it was. Our acceptable use policy bans sexual content involving minors. It bans using the service to attack systems you don’t own, or to target real people. Accounts that break it are closed. Taking the reflex out of the weights does not make those uses allowed. A person enforces the line, instead of a classifier guessing on every token.