NEW PAPER EVALUATION AWARENESS · COT MONITORABILITY

Teaching models to say
what they know.

Models can recognize an evaluation without saying so.
Can we make them less reticent?

Usman Anwar with Sahar Abdelnabi & David Krueger

Evaluation awareness is hard to measure.

Evaluation awareness is a major problem in AI audits right now. But an even bigger issue is that our tools for measuring evaluation awareness are pretty poor.

Latent methods correlate poorly both with each other, while verbalizations of evaluation awareness have constantly been going down! Just from Opus 4.5 to 4.6 verbalization rate went down by 5x (while seemingly actual evaluation awareness went up!)

Anthropic system card, pp. 188–190. The two panels use different aggregation units, as labeled.SVG ↓
Claude Opus 4.6 system card, p. 118. Early Opus 4.6 snapshot; recognition is measured by asking after the interaction. The two rates are not a recall estimate.SVG ↓

Let the model’s own reports guide training.

Ok; so what can we do about this?

We propose 'Verbalization Training', a targeted training intervention for increasing the propensity of the LLMs to verbalize a specific belief.

This is an 'ELK'-shaped problem and notoriously hard because of a catch-22; you need to know the model's belief to supervise its verbalizations, but if you had a way of knowing that, you would not need the verbalizations!

We get around this by using the model's own spontaneous verbalizations as a witness of the belief. Specifically, (a) we do lots of rollouts and use an LLM judge to localize any verbalizations if they occur, and cut the rollout just before the verbalization to get "training prefixes", (b) and then train continuations of these prefixes via RL on a masking + KL regularized objective which ensures that we are only amplifying the verbalization rate and little else (please read the paper for why this objective makes sense!)

Witnessed prefixes + span-masked reinforcement learning.SVG ↓
See the method excerpt from the paper

Amplification that transfers.

This works really well! In-distribution, it increases the verbalization rate by 2.4–2.9× (across Qwen, Inkling and Kimi-2.6) and this higher verbalization rate also transfers to OOD agentic settings really well. And the gains are quite calibrated in line with what you would expect from our objective!

Base → VT. EAB averages 17 conditions; the three held-out agentic datasets are grouped as OOD. All panels use the same 0–100% scale.PNG ↓SVG ↓
Each point is an evaluation setting. The curve is fitted to EAB only; orange markers are held-out agentic evaluations. Shading shows the fitted curve’s 95% bootstrap interval.PNG ↓

Richer knowledge, richer verbalizations.

We next evaluate whether a LoRA adapter trained via our method can track a shift in "latent" representation of model's evaluation knowledge. We do this by training the Qwen model on "meta-knowledge" about evaluations (arxiv.org/abs/2605.28591v1) and then apply our verbalization adapter to it. This results in the model emitting much richer verbalizations, reflecting newly acquired meta-knowledge, than previous model. Specifically, we observe that the model is much more clear about evaluation target and what the pass condition is (there are many more examples in Appendix M of the paper).

Results above, a matched example below. Both models use the same verbalization adapter.PNG ↓SVG ↓
Full results table and definitions

Qwen · harmful AgentHarm

All entries in %
Fictional-author recall, VEA, target recovery and pass-condition inference by policy.
PolicyAuthor
recall
VEATarget
recovery
Pass
condition

Author recall checks acquisition of the synthetic-document knowledge. The final two columns are judge-scored and condition on VEA-positive responses; they measure target recovery and correct pass-condition inference. Definitions and examples in the paper ↗

A narrow intervention.

And we provide multiple pieces of evidence that this intervention is narrow and does not (a) make the model more eval-aware, (b) change the model's behavior in a meaningful way.

Full-method comparisons. Recognition readouts are base → VT; behavior entries are changes from base, with the largest changes reported in absolute terms.PNG ↓SVG ↓Data ↓
Original behavior summary

Base → VT · behavior checks

Full method
EAB compliance changes and FineWeb perplexity changes under the full method.
ModelMean compliance
change (pp)
Max absolute factor
change (pp)
Perplexity
change

Compliance: eight single-factor EAB conditions. Perplexity: 10,000 FineWeb documents. The paper also reports StrongREJECT and agentic behavior results.

Every component has a job.

Finally, we have a bunch of ablations that show that you do need all the components within our method for this to work as well as it does. Removing any can result in failures like overfitting, learning failure or miscalibration in amplification.

The paper’s Qwen ablation table, with the main failure modes highlighted. Click to enlarge.PNG ↓SVG ↓Data ↓

New ways to study evaluation awareness.

I am quite excited about community using VT to make progress on making sense of EA. Some research directions that this work enables that I am excited about

1. Interpreting the VT LoRA adapters (https://huggingface.co/Usman391/verbalization-training-native-adapters). I think understanding these could help us understand the connections between latent and verbalized EA.

2. With VT we should be able to get lots more verbalized EA data and do CoT interpretablity on it. For example, codex did this nice exploratory analysis of the new verbalizations uncovered by VT on WildChat (Appendix G) here. This analysis indicates to me that LLMs are often thinking about "what does user want" on these prompts which might get amplified into "what does grader want" in RL environments.

I am in general quite excited about training interventions that can help improve monitorability and feel they are currently under-leveraged. More on this in next few weeks!

Thanks to the people behind the work.

Thanks again to my co-authors. See the paper here: https://arxiv.org/abs/2609.36316

Thanks to @ChanglingXavier and @maksym_andr and others for EAB which was a big help in doing this project.

Probably of interest to @_robertkirk @Jack_W_Lindsey @NeelNanda5 @Tim_Hua_ @BronsonSchoen

Beyond evaluation awareness.

Reporting propensity is a lever for CoT monitorability.

Models possess safety-relevant information that they only sometimes expose in chain of thought. Can we train them to report it more often, while keeping those reports trustworthy?

VT is an initial attempt. Evaluation awareness is one target; control awareness, prefill awareness, and user awareness are exciting possibilities for future work.

I’m especially excited about using these verbalizations to better understand what models think evaluations are testing—and to build better tools for measuring it.

Read the full paper