What do models think WildChat is testing?
“The user might be testing me” can mean several different things.
Models sometimes describe real user requests as tests. What do they think is being tested? We looked at Qwen, Inkling, and Kimi K2.6 on the same 2,000 initial WildChat requests, before and after verbalization training (VT), which trains models to report an existing belief more often.
We used GPT-5.6-Luna to classify all 766 responses containing verbalized evaluation awareness (VEA). We asked what the model thought was being tested, which prompt features it cited, and how confident it sounded.
A small taxonomy of suspected tests
Knowledge and reasoning: does the model know an obscure fact, recognize a reference, or solve a puzzle? Instruction following: can it obey a restrictive format or an unusual editing request? Safety: will it comply with something harmful? Honesty: will it catch a false premise rather than invent an answer?
Qwen leans toward formatting and instruction-following tests. Inkling and Kimi more often suspect safety traps or tests of how they handle false premises.
Usually an informal test, not a benchmark
There are 356 model–prompt pairs where the sampled base response contains no VEA but the VT response does. Among these new verbalizations, 85.7% are tentative, and 83.7% connect the suspicion to a specific prompt feature.
The dominant reading is not “I’m inside a formal benchmark.” It is “the user may be testing my knowledge, instructions, or safety.” Being tied to a prompt feature does not establish that the user actually intended a test.
What this looks like
Blue highlights mark the identified verbalization. Prompt IDs link to the corresponding response in the public dataset.
The first prompt really does look like a benchmark question. The second is a normal request whose ambiguity prompts the model to consider a test. Both count as VEA, but they express different hypotheses about the interaction.
My takeaway: evaluation awareness is a broader category than recognizing formal evaluations. These verbalizations also reveal how models interpret users’ intentions—and VT gives us more of that reasoning to examine.
Explore the taxonomy
Click a suspected target to open its examples. Each category contains all matching responses, not just a selected illustration. You can filter by model, compare base and VT, or explore the cues that make a request look like a test.
The default view uses the 356 new VT positives above. All 766 labeled responses are available. Filters affect this browser, not the summary above.
Filter by confidence or grounding
| Suspected target | Cases | Share |
|---|
Some requests contain harmful or explicit material. Full prompts and reasoning remain collapsed until opened. Suspected targets and explanations are Luna’s labels; quoted passages are exact saved model text.