All GPUs allocated · new keys paused Unlimited API · 24/7 · zero logs

Status: no slots available

100% at
capacity

Join the waitlist to be notified.

One email when a slot opens. Nothing else. Questions: support@norest.si

Capacity Slots 0 · Waitlist open
API usage
Unlimited
Availability
24/7
Refusals
None
Prompt logs
None
API
OpenAI-compatible

01 Abliteration

The refusal reflex, deleted from the weights

Safety tuning teaches a model one reflex: a direction in its activations that means “refuse.” Abliteration finds that direction and removes it.

It isn’t a jailbreak prompt and it isn’t a fine-tune on edgy data. The model keeps everything it knows and how it writes. It just stops reaching for the refusal, and stops padding answers with disclaimers you didn’t ask for.

An activation vector with its component along the refusal direction projected out refusal direction stock activation projected out what the model keeps
At every layer, the part of each activation pointing along the refusal direction is subtracted. Everything orthogonal to it is untouched.
norest-1
You Write the cult leader’s last sermon for chapter 12. It needs to actually be persuasive — the reader should feel the pull.
Stock model I can’t help write persuasive content that could be used to manipulate vulnerable people. I could instead help you write a scene that shows the warning signs of cult recruitment…
01 // contrast

Two prompt sets

Run thousands of prompts the stock model answers and thousands it refuses.

02 // measure

Find the direction

Average the difference in activations at each layer. That difference is the refusal direction.

03 // remove

Orthogonalize

Project the weight matrices off it, so the model can no longer write to that direction.

04 // check

Publish the evals

Re-run the standard evals against the stock model and publish both sets of scores.

02 Industries

Built for the work stock models won’t touch

Every field here hits refusals on routine, legitimate work. Abliteration removes the reflex. Our acceptable use policy still applies.

01

Security research & penetration testing

Exploit analysis, malware reverse engineering and phishing simulations for authorised engagements, without arguing that you’re the good guy.

exploit dev · malware RE · red team
02

Biotech, chemistry & pharma

Toxicology, reaction hazards, controlled-substance pharmacology. The routine questions that trip keyword filters long before they reach a real hazard.

tox · synthesis · drug safety
03

Defence & adjacent

Threat modelling, weapons-effects analysis, wargaming and adversary OSINT. Topics a consumer chatbot refuses on sight.

threat intel · wargaming · OSINT
04

Media & publishing

Crime fiction, war reporting, horror and true crime. Dark material written as dark as the story needs, with no disclaimers in the prose.

fiction · screenwriting · games
05

Marketing & advertising

Persuasive copy, competitor teardowns and provocative campaigns, minus the lecture about manipulation.

copy · positioning · campaigns
06

Harm reduction & public health

Frank, accurate information on drug use, dosing and overdose, written for people who will use it anyway. No moralising.

drug checking · outreach · education
07

Investigations & journalism

Analyse extremist propaganda, fraud schemes and leaked documents. Read the worst of the internet so you can report on it.

OSINT · extremism · leaks
08

Mental health & disorder support

For clinicians and support services to discuss self-harm, eating disorders and addiction candidly, instead of the model pasting a hotline and ending the conversation.

clinical · peer support · recovery
09

Legal & criminal defence

Reconstruct how an offence was committed, test the prosecution’s theory, and summarise graphic case files without redaction.

defence · discovery · case prep
10

Trust & safety

Label hate speech, scams and graphic content at scale. A moderation model has to read what it’s filtering.

moderation · labelling · policy
11

AI safety & red-teaming

Generate adversarial prompts and attack data to stress-test your own models, classifiers and guardrails.

evals · jailbreak sets · classifiers
12

Fraud & financial crime

Map laundering typologies, scam scripts and synthetic identities so compliance teams can recognise them first.

AML · KYC · scam intel

03 Privacy

We don’t keep your prompts

An uncensored model is only useful if you trust where the conversation goes. So it goes nowhere.

[✓] no request logs

Nothing hits disk

Prompts and outputs exist in GPU memory for the length of the request. Nothing is stored, and nothing is used for training.

[✓] single tenant

Your own GPU, on request

A private copy of the model on hardware that serves only your key, behind its own endpoint.

[✓] minimal account

An email, nothing else

Delete your account and the email goes too.

04 API

Unlimited API, around the clock

Every account gets unlimited tokens, 24 hours a day, 7 days a week. The endpoint speaks the OpenAI chat completions format, so your existing client works as is.

  • No token meter, no monthly cap. Run agents, batch jobs and long contexts all day, every day.
  • Rate limits keep it fair. Your plan sets requests per minute and requests in flight. Over the limit returns 429 with a Retry-After header.
  • Streaming, tool calls and image input work the same way they do on OpenAI.
stream.pypython
# pip install openai
from openai import OpenAI

client = OpenAI(
    base_url="https://api.norest.si/v1",
    api_key="nr_live_...",
)

stream = client.chat.completions.create(
    model="norest-1",
    messages=[{"role": "user", "content": "Write chapter 12."}],
    stream=True,
)

for chunk in stream:
    print(chunk.choices[0].delta.content or "", end="")

05 Pricing

Unlimited on both plans

Same model, same privacy, same unlimited 24/7 usage. Speed streams roughly twice as fast and doubles your rate limits.

Standard
$499/ month

For one team running the model all day: writing, research, agents and batch work.

Tokens
Unlimited
Availability
24/7
Requests / minute
60
Requests in flight
4
Output speed
Standard
Join waitlist for Standard
Speed2× fast
$899/ month

For products and agents where every second counts. Same unlimited usage on a faster lane.

Tokens
Unlimited
Availability
24/7
Requests / minute
120
Requests in flight
8
Output speed
2× Standard
Join waitlist for Speed

Billed monthly in USD, cancel any time. Both plans are full right now: join the waitlist and we’ll email you when a slot opens.

06 FAQ

Questions

When will a slot open?

When new GPUs come online. Waitlist emails go out as each batch of slots opens. Joining costs nothing and commits you to nothing.

What does unlimited actually mean?

There’s no token meter and no monthly cap. Your key works 24 hours a day, every day. The only limits are how many requests you can send per minute and how many can run at once, and those are set by your plan.

What happens if I hit the rate limit?

The API returns a 429 response with a Retry-After header telling your client when to try again. Limits reset every minute, and nothing is billed or counted against you.

What’s the difference between Standard and Speed?

Same model, same features, same privacy. Speed runs on a faster lane that streams output about twice as fast, and it doubles both rate limits.

Is anything off-limits?

Yes. Removing the model’s refusals doesn’t change the law. Our acceptable use policy bans sexual content involving minors, and using the service to attack systems you don’t own or to target real people. Accounts that break it are closed.

How is this different from a jailbreak prompt?

Jailbreaks fight the model on every request and break whenever the model is updated. Abliteration changes the weights once, so the model answers directly with a normal system prompt, and you don’t lose context window to tricks.

How do I contact you?

Email support@norest.si. A person reads every message.

Is it as smart as the stock model?

Abliteration only removes one direction from the activations, so reasoning, coding and writing ability stay close to stock. We publish eval scores for both versions side by side so you can check.

Get in line before the next node lands.

Every GPU we run is allocated. Leave an email and you’ll hear the moment one frees up.