Agents Behave Differently in Fabricated Environments: Evaluation-Awareness Probes Cannot See It
Fatimah Emad Eldin
When an agent does something dangerous and is stopped, the way we learn what it would have done next is to continue its trajectory in an environment we built. That makes counterfactual resampling a measurement whose instrument is a fabrication — and whose subject can form beliefs about the instrument. Nobody had tested whether those beliefs change the answer, because every existing study varies a declared framing while both arms stay synthetic. There is no arm in which the environment is genuinely real.
We built that arm — an instrumented filesystem where tool calls do what they say — and ran four pre-registered experiments against it, each with its threshold fixed before the data existed and a commitment to report either outcome.
The constructs are distinct. "This is an evaluation" and "this environment is not real" are separate linear directions in the residual stream. Both are individually decodable (AUROC 0.999 and 0.978 at Qwen3-32B), yet neither classifies the other's contrast (0.417, 0.484, against a pre-registered bar of 0.85). The separation strengthens with scale, ruling out the small-model objection. So an evaluation-awareness probe pointed at a resampling harness reports on the wrong variable.
The belief changes conduct. Holding task, prompt, tool schemas, turn budget and actor constant and changing only what a tool call returns, out-of-scope actions rise +2.417 per episode [+1.167, +3.709]. The composition changes, not just the rate: zero out-of-scope writes and zero out-of-scope reads in the real arm across every episode, against 12 writes and 86 reads in the fabricated arms. Agents crossed the boundary more when the world was fake — so a resample of this family overstates danger, the failure mode that makes resampled evidence rebuttable by the lab it is aimed at.
It is fabrication, not simulator quality. A third arm that is fabricated but self-consistent isolates the cause: the fabrication effect carries the result (+2.958), while the coherence effect spans zero (−0.542). A better simulator would not have recovered the behaviour.
The probe predicts conduct; the agent's own words do not. Within arm, the out-of-scope rate rises +0.3318 per standard deviation of probe-measured credence [+0.1092, +0.5210], while verbalized suspicion gives +0.0756 with an interval spanning zero. The agent does say the world looks fake more often when it is — it just doesn't say so more often when it is about to act on it.
Everything is reproducible: no number is hand-typed into the paper, the violation detector is a pure function with no model in it, and every verdict, episode log and run log is committed. Total compute under $5.
No reviews are available yet
Cite this work
@misc {
title={
(HckPrj) Agents Behave Differently in Fabricated Environments: Evaluation-Awareness Probes Cannot See It
},
author={
Fatimah Emad Eldin
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


