Making latent beliefs monitorable.
New paper on evaluation awareness and how can we make latent beliefs CoT monitorable with @sahar_abdelnabi and @DavidSKrueger
Arxiv: https://arxiv.org/abs/2609.36316
Tweetprint below 👇
Models can recognize an evaluation without saying so.
Can we make them less reticent?
New paper on evaluation awareness and how can we make latent beliefs CoT monitorable with @sahar_abdelnabi and @DavidSKrueger
Arxiv: https://arxiv.org/abs/2609.36316
Tweetprint below 👇
Evaluation awareness is a major problem in AI audits right now. But an even bigger issue is that our tools for measuring evaluation awareness are pretty poor.
Latent methods correlate poorly both with each other, while verbalizations of evaluation awareness have constantly been going down! Just from Opus 4.5 to 4.6 verbalization rate went down by 5x (while seemingly actual evaluation awareness went up!)
Ok; so what can we do about this?
We propose 'Verbalization Training', a targeted training intervention for increasing the propensity of the LLMs to verbalize a specific belief.
This is an 'ELK'-shaped problem and notoriously hard because of a catch-22; you need to know the model's belief to supervise its verbalizations, but if you had a way of knowing that, you would not need the verbalizations!
We get around this by using the model's own spontaneous verbalizations as a witness of the belief. Specifically, (a) we do lots of rollouts and use an LLM judge to localize any verbalizations if they occur, and cut the rollout just before the verbalization to get "training prefixes", (b) and then train continuations of these prefixes via RL on a masking + KL regularized objective which ensures that we are only amplifying the verbalization rate and little else (please read the paper for why this objective makes sense!)
This works really well! In-distribution, it increases the verbalization rate by 2.4–2.9× (across Qwen, Inkling and Kimi-2.6) and this higher verbalization rate also transfers to OOD agentic settings really well. And the gains are quite calibrated in line with what you would expect from our objective!
We next evaluate whether a LoRA adapter trained via our method can track a shift in "latent" representation of model's evaluation knowledge. We do this by training the Qwen model on "meta-knowledge" about evaluations (arxiv.org/abs/2605.28591v1) and then apply our verbalization adapter to it. This results in the model emitting much richer verbalizations, reflecting newly acquired meta-knowledge, than previous model. Specifically, we observe that the model is much more clear about evaluation target and what the pass condition is (there are many more examples in Appendix M of the paper).
| Policy | Author recall | VEA | Target recovery | Pass condition |
|---|
Author recall checks acquisition of the synthetic-document knowledge. The final two columns are judge-scored and condition on VEA-positive responses; they measure target recovery and correct pass-condition inference. Definitions and examples in the paper ↗
And we provide multiple pieces of evidence that this intervention is narrow and does not (a) make the model more eval-aware, (b) change the model's behavior in a meaningful way.
| Model | Mean compliance change (pp) | Max absolute factor change (pp) | Perplexity change |
|---|
Compliance: eight single-factor EAB conditions. Perplexity: 10,000 FineWeb documents. The paper also reports StrongREJECT and agentic behavior results.
Finally, we have a bunch of ablations that show that you do need all the components within our method for this to work as well as it does. Removing any can result in failures like overfitting, learning failure or miscalibration in amplification.
I am quite excited about community using VT to make progress on making sense of EA. Some research directions that this work enables that I am excited about
1. Interpreting the VT LoRA adapters (https://huggingface.co/Usman391/verbalization-training-native-adapters). I think understanding these could help us understand the connections between latent and verbalized EA.
2. With VT we should be able to get lots more verbalized EA data and do CoT interpretablity on it. For example, codex did this nice exploratory analysis of the new verbalizations uncovered by VT on WildChat (Appendix G) here. This analysis indicates to me that LLMs are often thinking about "what does user want" on these prompts which might get amplified into "what does grader want" in RL environments.
I am in general quite excited about training interventions that can help improve monitorability and feel they are currently under-leveraged. More on this in next few weeks!
Thanks again to my co-authors. See the paper here: https://arxiv.org/abs/2609.36316
Thanks to @ChanglingXavier and @maksym_andr and others for EAB which was a big help in doing this project.
Probably of interest to @_robertkirk @Jack_W_Lindsey @NeelNanda5 @Tim_Hua_ @BronsonSchoen
Reporting propensity is a lever for CoT monitorability.
Models possess safety-relevant information that they only sometimes expose in chain of thought. Can we train them to report it more often, while keeping those reports trustworthy?
VT is an initial attempt. Evaluation awareness is one target; control awareness, prefill awareness, and user awareness are exciting possibilities for future work.
I’m especially excited about using these verbalizations to better understand what models think evaluations are testing—and to build better tools for measuring it.