Why counting refusals can miss how a model steers an answer
A research preprint separates detecting a sensitive concept from choosing a response. It suggests a broader way to read claims about uncensored AI.
Intelligence belongs
in the open.
Refusals, alignment and what evaluations can actually tell us.
2 stories
A research preprint separates detecting a sensitive concept from choosing a response. It suggests a broader way to read claims about uncensored AI.
A July research preprint examines changes in decision behavior after abliteration. Its narrow experiment raises a wider question about model modifications.