Honest vs. Deceptive Feedback: How an Overseer’s Truthfulness Affects a Language Model’s Task Success and Internal Valence
Itbaan Safwan, Muhammad Shayan Shamsi
We test whether an overseer's honesty shapes both a language model's performance and its internal valence, without ever using emotional language in the prompts. A Qwen3-4B "student" solves hard mazes over many turns while a frontier "overseer" gives either honest (teacher) or covertly deceptive (adversary) feedback, and we read the student's internal valence each turn by projecting its activations onto a pre-identified welfare direction. Across 30 paired mazes, deception roughly halves the solve rate (63% → 30%; McNemar p ≈ 0.021) and reliably lowers the internal valence signal (paired t(29) = 4.26). Since no evaluative language appears anywhere in the text, the signal reflects the quality of the interaction the model is placed in rather than any scripted emotional performance, suggesting internal valence probes may capture something about how a model is being treated.
This study builds on work on the “functional welfare axis” reported in Han et al. (2026). Its research question is two-fold: on the one hand, it concerns to what extent internal activity patterns along this axis vary when success at a maze-solving task is high/low; on the other hand, it is about investigating whether deceptive vs. honest feedback on a maze-solving task influences the model’s success at the task.
The results are: the type of feedback has an influence on task performance (worse performance in the deceptive condition) and the internal valence signal is reliably lower in the deception condition.
The paper suggests that “the trustworthiness of feedback, not task difficulty alone, shapes a subject’s experience of a task.” (p. 2), and “The signal tracks the honesty of guidance the model is never told about, suggesting internal valence probes may capture something about how a model is being treated” (p. 5).
From reading the paper, it’s not clear to me that this interpretation is warranted. If I understand correctly, the only feedback received by the model is the overseer’s feedback (so no explicit “maze solved” / “maze not solved” feedback). But it’s not clear from the presentation to what extent the overseer’s feedback entails feedback on whether a maze is solved or not. That is, instead of tracking honesty or trustworthiness, the valence signal could also track success at the maze solving task. I would recommend addressing this confound explicitly (or saying how your design addresses it).
This is a strong, well-motivated sprint project that introduces a ground-truth-verifiable, multi-turn setup for studying how honest versus deceptive oversight relates to task performance and an internal valence probe. The paired design, independent environment scoring, appropriate statistical testing, and candid discussion of limitations are notable strengths. The main limitation is causal attribution: the two feedback conditions differ in semantics, induced confidence, and task progress, while the welfare direction may also track truth, assent, confidence, or goal achievement. The experiment therefore establishes an association between feedback condition, performance, and probe activation more convincingly than it isolates deception or model welfare. Matched-text controls, feedback-only baselines, repeated trajectories, and additional models would substantially strengthen the conclusion. Overall, this is focused, technically competent, and promising hackathon work.
Cite this work
@misc {
title={
(HckPrj) Honest vs. Deceptive Feedback: How an Overseer’s Truthfulness Affects a Language Model’s Task Success and Internal Valence
},
author={
Itbaan Safwan, Muhammad Shayan Shamsi
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


