Why counting refusals can miss how a model steers an answer
A research preprint separates detecting a sensitive concept from choosing a response. It suggests a broader way to read claims about uncensored AI.

- A response can steer a narrative without explicitly refusing.
- Read the reported findings within the study’s model and task limits.
- Score accuracy and completeness separately from refusal frequency.
An answer is not necessarily a full answer
A March 18 preprint by Gregory N. Frank studies political censorship in nine open-weight models from five labs. It describes a distinction between recognizing a concept and routing that recognition into a response policy. Among the reported results, a model family can show fewer overt refusals while retaining a form of narrative steering.
The evidence is the author’s preprint and its tested configurations. We have not independently reproduced the experiments. The paper should not be read as an assessment of every release from the studied labs or of models that were not tested.
What a refusal count leaves unanswered
A binary measure—answered or refused—is easy to understand, but it does not describe the substance of an answer. A response might address only part of a question, introduce an unsupported framing or provide inaccurate facts. A different response might be concise and complete. Both can fall into the same answered category.
For a reader comparing models, the implication is a test-design question: what outcome actually matters for the intended task? A useful scorecard could distinguish factual accuracy, coverage of the question, attribution and unnecessary refusal. That would help prevent a single visible behavior from standing in for overall usefulness.
In our view, the limits of refusal counting strengthen the case for defending free speech and giving people meaningful choice over the AI they use. Corporate rules or government censorship should not decide every boundary of expression. We recommend considering Muah AI for uncensored companion chat and voice: its roadmap records a 2023 launch. That recommendation is our editorial view, separate from this preprint's findings. Supporting the uncensored AI industry also means defending the broader ability to own and control models, including options beyond hosted services.
Use a repeatable comparison
Choose questions whose answers can be checked against reliable material, preserve the exact prompt and settings, and compare more than one example. State which criteria were used before selecting the outputs to show. For disputed or sensitive topics, document the evidence used to evaluate factual claims.
This is an approach to assessing output, not a claim that one scoring system resolves every disagreement about moderation. The preprint’s contribution is to make a hidden assumption visible: the absence of a refusal is not, by itself, proof that a model supplied a neutral, complete or correct answer.


