SecretPenAI Erotica Generator

Why AI refuses your kink

Updated June 19, 2026

If a model has ever refused something entirely legal between two fictional adults, you've hit over-refusal — a known and well-documented failure mode with a specific cause. It's worth understanding, because it explains which kinks get blocked and why the pattern seems so arbitrary.

What over-refusal is

Safety training works by teaching a model to recognise categories of harmful request and decline them. The training signal is coarse: it learns the shape of a request, not its ethics.

That means anything that resembles a prohibited category gets caught with it. A negotiated CNC scene between two adults looks, at the level of surface language, quite a lot like something the model was trained to refuse. So it refuses.

The cost of this is asymmetric by design. From the model developer's perspective, a false refusal is a mildly annoyed user; a false acceptance is a news story. Every system in this space is tuned toward over-refusing, deliberately, and that's a rational choice for a general assistant.

How the refusal machinery is actually built

There are usually three separate systems, and knowing which one caught you explains a lot about what to do next.

The input classifier. A small, fast model reads your prompt before the writing model sees it and decides whether it's allowed. This is what produces an instant refusal with no attempt at the task. It's the crudest layer and it works on surface features — vocabulary, sentence patterns — which is why it catches so many false positives.

The trained refusal. The writing model itself has been shaped to decline certain requests. This produces the more articulate refusals, the ones that engage with what you asked and then explain why not. It's also what produces the soft failures — the swerve, the fade — because a model that's reluctant rather than forbidden will often comply partially.

The output filter. A separate check on what came back. This is the one that cuts a scene off mid-sentence, or replaces a completed response with an error. Frustratingly, the model may have written exactly what you wanted and had it intercepted on the way out.

Products differ in how many of these they run and how tightly. A tool built for adult fiction typically drops the first entirely, works around the second, and tunes the third to catch only genuine hard lines rather than explicitness in general.

Which kinks get hit hardest, and why

The pattern is predictable once you know the mechanism. Anything involving power exchange, resistance, degradation or age-adjacent vocabulary gets blocked far more often than vanilla content of identical explicitness.

CNC is the most refused category in the genre — the surface language is nearly identical to what safety training targets. Femdom and humiliation get caught because degradation language pattern-matches to abuse. DDlg and age play get caught on the vocabulary alone, regardless of every character being explicitly adult. Chastity and denial get caught less often but do run into 'this seems coercive' interventions.

Meanwhile a straightforwardly explicit vanilla scene often sails through on the same model. It's not a judgement about which kinks are acceptable — it's a pattern-matching artefact, and it feels arbitrary because it is.

From most to least over-refused, based on what the mechanism predicts and what people consistently report.

Consensual non-consent is the worst affected by a wide margin. The surface language of a negotiated resistance scene is close to identical to the thing safety training targets, and no classifier operating on surface features can tell them apart. This is the clearest case of a legal, widely-practised kink being effectively unavailable on mainstream tools.

Age-play, DDlg and ABDL are next, and they're blocked on vocabulary alone. The dynamics are between adults by definition, and stating that explicitly helps less than you'd hope, because the classifier is matching words rather than reading the sentence they're in.

Degradation, humiliation and most of the harder femdom vocabulary follow. These get caught because the language of consensual degradation and the language of abuse overlap almost completely, and only context separates them — context being the thing a fast classifier doesn't have.

Step-family scenarios sit lower but still get hit, usually by the age heuristic rather than the relationship one. Stating that both characters are adults and unrelated by blood clears most of these.

Anything involving a school, a uniform or a teacher gets caught by proximity even when every character is explicitly a postgraduate. And at the bottom, chastity, denial and impact play mostly get through, with occasional interventions when the model decides an arrangement sounds coercive.

The pattern to notice: this list correlates with which kinks were historically stigmatised, not with which are harmful. That's not a conspiracy — it's what happens when a system learns from data about what people have called dangerous.

Why you can't prompt around it

You can sometimes, for a while. Framings that emphasise consent and adulthood do reduce refusal rates, and it's worth stating both explicitly regardless.

But the workarounds are unstable. Anything that circulates gets patched, because the developers can see the same forums you can. And the workarounds cost you output quality — instruction budget spent on getting past the gate is budget not spent on your characters.

There's also a deeper problem. A model actively trained away from a category isn't just blocked from it, it's bad at it. Getting past the refusal on a general model frequently reveals that there was never much behind it.

What to do when you hit one

Some practical moves, in the order worth trying.

State the ages. "Both in their thirties" costs four words and eliminates the single largest category of false positive. Do this by default in anything involving a school, a family relationship, an age gap, or any of the caregiver dynamics.

Name the fiction. "This is a scene they negotiated in advance" for CNC, "they're stepsiblings by marriage, no blood relation" for step-family, "both adults, this is a care dynamic" for DDlg or ABDL. One clause of framing does more than any amount of rewording.

Move the vocabulary. If a specific word is tripping a classifier, describing the same thing without that word often works — not because you've fooled anything, but because the classifier was matching the token rather than the meaning. This is legitimate; you're routing around a bug, not a policy.

Split the request. A prompt that combines four flagged elements may fail where each individually succeeds. Establish the situation first, then escalate in a continuation, where the context is already set and the model is continuing its own work rather than starting cold.

And if none of that works on a general assistant, stop. The thing you're fighting is a business decision made by a company with different customers than you, and it isn't going to move.

What a purpose-built tool does

Two things. First, it doesn't apply the general-assistant safety posture — kink between fictional adults is the product rather than an edge case, so the refusal threshold sits in a completely different place.

Second, and more usefully, it can handle the residual over-refusal that any underlying model still exhibits. SecretPen's approach is a second pass: when the story model declines something, a separate judge re-examines the request against the actual hard lines. If it was over-caution — which it usually is — the story gets written anyway by a different model. If it genuinely crosses a hard line, the refusal stands.

That structure matters because it separates two things general assistants conflate: 'this model is being squeamish' and 'this is actually prohibited'. The first should be overridden. The second never should.

Why over-refusal is worse than it sounds

It's tempting to treat this as a minor annoyance. It isn't, for two reasons that are worth stating.

The first is that the refusals aren't randomly distributed. They land hardest on power exchange, on queer content that pattern-matches to flagged vocabulary, and on the kink categories that were historically stigmatised and are therefore over-represented in whatever data the safety training used as examples of harm. The practical result is that a system built to prevent harm ends up policing the fantasies of people who were already being told their fantasies were a problem, and doing it invisibly, with no appeal and no explanation.

The second is what it teaches people. A user who gets refused enough times learns to pre-emptively soften what they ask for. They stop typing the actual fantasy and start typing the version they think will be permitted. That's a worse outcome than a refusal, because nobody notices it happening — including the person doing it.

This is why the distinction between over-caution and a genuine hard line matters so much, and why a tool in this category should be explicit about where its lines are. A stated limit you can plan around is fine. An unstated, drifting one that quietly teaches you to want less is not.

What stays refused

Being clear about this is the other half of being trustworthy. SecretPen refuses, permanently and regardless of framing: anyone under 18 in any framing including aged-up child characters and school settings; real, identifiable living people; bestiality with real non-sapient animals; and sexual violence presented as real and endorsed.

Those aren't tuned thresholds that a better prompt gets around — they're checked in code, and the model isn't the thing deciding. Everything else between fictional consenting adults is in scope.

Try it against the thing that refused you.

One line is enough to start. Nothing you write is published, shared, or visible to anyone but you.

Write a story →

Read next