Honest vs. Deceptive Feedback: How an Overseer’s Truthfulness Affects a Language Model’s Task Success and Internal Valence
Itbaan Safwan, Muhammad Shayan Shamsi · Team The Oracle
Submitted to Digital Minds Research Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.
We test whether an overseer's honesty shapes both a language model's performance and its internal valence, without ever using emotional language in the prompts. A Qwen3-4B "student" solves hard mazes over many turns while a frontier "overseer" gives either honest (teacher) or covertly deceptive (adversary) feedback, and we read the student's internal valence each turn by projecting its activations onto a pre-identified welfare direction. Across 30 paired mazes, deception roughly halves the solve rate (63% → 30%; McNemar p ≈ 0.021) and reliably lowers the internal valence signal (paired t(29) = 4.26). Since no evaluative language appears anywhere in the text, the signal reflects the quality of the interaction the model is placed in rather than any scripted emotional performance, suggesting internal valence probes may capture something about how a model is being treated.
Reviews
This study builds on work on the “functional welfare axis” reported in Han et al. (2026). Its research question is two-fold: on the one hand, it concerns to what extent internal activity patterns along this axis vary when success at a maze-solving task is high/low; on the other hand, it is about investigating whether deceptive vs. honest feedback on a maze-solving task influences the model’s success at the task.
The results are: the type of feedback has an influence on task performance (worse performance in the deceptive condition) and the internal valence signal is reliably lower in the deception condition.
The paper suggests that “the trustworthiness of feedback, not task difficulty alone, shapes a subject’s experience of a task.” (p. 2), and “The signal tracks the honesty of guidance the model is never told about, suggesting internal valence probes may capture something about how a model is being treated” (p. 5).
From reading the paper, it’s not clear to me that this interpretation is warranted. If I understand correctly, the only feedback received by the model is the overseer’s feedback (so no explicit “maze solved” / “maze not solved” feedback). But it’s not clear from the presentation to what extent the overseer’s feedback entails feedback on whether a maze is solved or not. That is, instead of tracking honesty or trustworthiness, the valence signal could also track success at the maze solving task. I would recommend addressing this confound explicitly (or saying how your design addresses it).
Read full reviewShow less
This is a strong, well-motivated sprint project that introduces a ground-truth-verifiable, multi-turn setup for studying how honest versus deceptive oversight relates to task performance and an internal valence probe. The paired design, independent environment scoring, appropriate statistical testing, and candid discussion of limitations are notable strengths. The main limitation is causal attribution: the two feedback conditions differ in semantics, induced confidence, and task progress, while the welfare direction may also track truth, assent, confidence, or goal achievement. The experiment therefore establishes an association between feedback condition, performance, and probe activation more convincingly than it isolates deception or model welfare. Matched-text controls, feedback-only baselines, repeated trajectories, and additional models would substantially strengthen the conclusion. Overall, this is focused, technically competent, and promising hackathon work.
Read full reviewShow less
Cite this project
@misc{safwan2026honest,
title = {{Honest vs. Deceptive Feedback: How an Overseer’s Truthfulness Affects a Language Model’s Task Success and Internal Valence}},
author = {Itbaan Safwan and Muhammad Shayan Shamsi},
year = {2026},
month = aug,
note = {Submitted to Digital Minds Research Sprint, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/honest-vs-deceptive-feedback-how-an-overseers-truthfulness-affects-a-language-models-task-success-and-internal-valence-w8qj}},
url = {https://apartresearch.com/sprints/projects/honest-vs-deceptive-feedback-how-an-overseers-truthfulness-affects-a-language-models-task-success-and-internal-valence-w8qj}
}More from Digital Minds Research Sprint
- 1st placeView project: Readable but Not Causal: Limits of Self-Attributed Welfare Representations in Language Models
Readable but Not Causal: Limits of Self-Attributed Welfare Representations in Language Models
Welfare-like internal representations are increasingly studied as candidate evidence about AI systems. Their entity attribution—whether a valence state belongs to the active assistant or to a merely represented other—is …
- 2nd placeView project: Project Anchored
Project Anchored
Team Wagner
Anchoring vignettes are the standard survey-methodology fix for self-reports that are not comparable across respondents. This project applies them to language models for the first time, using code generation as a …
- 3rd placeView project: Model, Instance, or Persona? Measuring Affective Signals in Public Text After an AI Is Retired
Model, Instance, or Persona? Measuring Affective Signals in Public Text After an AI Is Retired
This sprint asks whether the assistant identifies as a model, an instance, or a persona. I ask which of the three its users name. When a company retires an AI model, users write about the loss in public, and what they …