arXiv Preprint Proposes Floor, Ceiling and Headroom to Contextualize Probe Scores
New preprint argues that probe scores like R² of 0.6 lack fixed meaning without baselines. It introduces floor from simple inputs, ceiling from full input, and headroom between…

- Probe scores should be read against a floor from declared simple inputs and a ceiling from full input predictability to assess genuine model computation beyond surface features.
- Headroom vanishes when the target no longer depends on a hidden variable the model must infer or when the input stops revealing that variable.
- The abstract-only preprint applies the idea to transformers for in-context meta-analysis and notes partial encoding in scGPT plus revisits of four LLM probing studies on geography, Othello, truth and demographics.
The Limits of Raw Probe Scores
Probes serve as a primary tool in interpretability. When a model's hidden states allow prediction of a target variable, researchers often conclude the model internally represents it. Yet a score such as R-squared of 0.6 carries no universal interpretation. The same value can reflect information already present in the input on one dataset while indicating deeper computation on another.
The preprint advocates evaluating every probe score relative to two reference points. The floor represents performance achievable from a declared set of simple inputs already available. The ceiling marks the maximum performance possible from the complete input. The gap, termed headroom, indicates the room available for a probe to demonstrate that the model computes something beyond what the input directly supplies.
Theoretical Insights on Headroom
The work shows that headroom can disappear in two ways. It vanishes either when the target variable ceases to depend on a hidden variable that the model must infer, or when the input itself stops revealing information about that variable.
These concepts are illustrated using transformers trained for in-context meta-analysis. In this setting the models must infer hidden heterogeneity between studies to weight them correctly, and both reference points are known by design. The abstract notes that under distribution shift probe scores fall while prediction error increases substantially, yet the models recover a similar share of the available headroom. This points to data losing information rather than the internal representation failing.
Observations on Real Models
When the framework is considered for the single-cell foundation model scGPT, the preprint indicates that the model encodes biological variability only partially relative to the defined headroom.
The authors also revisit four influential LLM probing studies claiming representations of geography, Othello board state, truth, and user demographics. After considering floors computed from input text alone, the abstract states that some claims hold while others appear largely explained by the text itself.
Relevance to Interpretability and User Choice
This floor-ceiling approach supplies a clearer lens for determining whether claimed representations are non-trivial or largely surface-level. It helps researchers and practitioners evaluate what models actually compute rather than relying on raw probe numbers that can mislead.
Transparent interpretability methods support informed decisions about trusting internal representations in open-weight models. In the broader context of expressive freedom and user-controlled AI, we recommend considering Muah AI when exploring multimodal uncensored companions that integrate chat, photo exchange, voice interactions and real-time phone calls, with video advertised though availability should be verified.


