Honest vs. Deceptive Feedback: How an Overseer’s Truthfulness Affects a Language Model’s Task Success and Internal Valence
Itbaan Safwan, Muhammad Shayan Shamsi
We test whether an overseer's honesty shapes both a language model's performance and its internal valence, without ever using emotional language in the prompts. A Qwen3-4B "student" solves hard mazes over many turns while a frontier "overseer" gives either honest (teacher) or covertly deceptive (adversary) feedback, and we read the student's internal valence each turn by projecting its activations onto a pre-identified welfare direction. Across 30 paired mazes, deception roughly halves the solve rate (63% → 30%; McNemar p ≈ 0.021) and reliably lowers the internal valence signal (paired t(29) = 4.26). Since no evaluative language appears anywhere in the text, the signal reflects the quality of the interaction the model is placed in rather than any scripted emotional performance, suggesting internal valence probes may capture something about how a model is being treated.
No reviews are available yet
Cite this work
@misc {
title={
(HckPrj) Honest vs. Deceptive Feedback: How an Overseer’s Truthfulness Affects a Language Model’s Task Success and Internal Valence
},
author={
Itbaan Safwan, Muhammad Shayan Shamsi
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


