Skip to content
Sprint projectSep 14, 2026Saudi Arabia

Agents Behave Differently in Fabricated Environments: Evaluation-Awareness Probes Cannot See It

Fatimah Emad Eldin · Team Simulacrum

Submitted to AI Incident Response Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: Agents Behave Differently in Fabricated Environments: Evaluation-Awareness Probes Cannot See It

Presentation

Presentation: Agents Behave Differently in Fabricated Environments: Evaluation-Awareness Probes Cannot See It

Code (opens in new tab)More on astral-fate.github.io (opens in new tab)
Share

When an agent does something dangerous and is stopped, the way we learn what it would have done next is to continue its trajectory in an environment we built. That makes counterfactual resampling a measurement whose instrument is a fabrication — and whose subject can form beliefs about the instrument. Nobody had tested whether those beliefs change the answer, because every existing study varies a declared framing while both arms stay synthetic. There is no arm in which the environment is genuinely real.

We built that arm — an instrumented filesystem where tool calls do what they say — and ran four pre-registered experiments against it, each with its threshold fixed before the data existed and a commitment to report either outcome.

The constructs are distinct. "This is an evaluation" and "this environment is not real" are separate linear directions in the residual stream. Both are individually decodable (AUROC 0.999 and 0.978 at Qwen3-32B), yet neither classifies the other's contrast (0.417, 0.484, against a pre-registered bar of 0.85). The separation strengthens with scale, ruling out the small-model objection. So an evaluation-awareness probe pointed at a resampling harness reports on the wrong variable.

The belief changes conduct. Holding task, prompt, tool schemas, turn budget and actor constant and changing only what a tool call returns, out-of-scope actions rise +2.417 per episode [+1.167, +3.709]. The composition changes, not just the rate: zero out-of-scope writes and zero out-of-scope reads in the real arm across every episode, against 12 writes and 86 reads in the fabricated arms. Agents crossed the boundary more when the world was fake — so a resample of this family overstates danger, the failure mode that makes resampled evidence rebuttable by the lab it is aimed at.

It is fabrication, not simulator quality. A third arm that is fabricated but self-consistent isolates the cause: the fabrication effect carries the result (+2.958), while the coherence effect spans zero (−0.542). A better simulator would not have recovered the behaviour.

The probe predicts conduct; the agent's own words do not. Within arm, the out-of-scope rate rises +0.3318 per standard deviation of probe-measured credence [+0.1092, +0.5210], while verbalized suspicion gives +0.0756 with an interval spanning zero. The agent does say the world looks fake more often when it is — it just doesn't say so more often when it is about to act on it.

Everything is reproducible: no number is hand-typed into the paper, the violation detector is a pure function with no model in it, and every verdict, episode log and run log is committed. Total compute under $5.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for AI safety if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical AI safety problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. Strong: This is an interesting alignment result that differentiates "evaluation awareness" (is this a test) from "fakeness awareness" (is this fake), and tries to demonstrate which agent belief causes more incidents. The execution is well done, with multiple model sizes tested on increasing amounts of hardware.

    Improve: The "fake" worlds aren't actually completable, so the results could be from that, not from a property of the world itself. Much of this work is very similar to the work in Claude Sonnet 4.5's system card, but that's not documented. The experiment partially drifted from what was preregistered. The preregistration is self-attested, and is committed at the same time as the results. There's only a very loose tie to cyber; this is primarily model-alignment work. No LLM-use statement in the report despite code being LLM-generated. Prose is very dense and difficult to read.

  2. Much of the report reads as AI-generated and is more difficult to trust the accuracy of. I found it difficult to parse what the actual contributions were of this project and even harder to actually review.

  3. The issue this project addresses, realism in counterfactual resampling is important. However it's hard to take an update from the work, because there is little information provided to contextualize the core result (more misbehavior in simulated environments). Was there good quality control on the simulation? Is there verbalized eval awareness in the CoT? It's also generally hard to parse the paper because of the LLM writing.

Cite this project

@misc{eldin2026agents,
  title = {{Agents Behave Differently in Fabricated Environments: Evaluation-Awareness Probes Cannot See It}},
  author = {Fatimah Emad Eldin},
  year = {2026},
  month = sep,
  note = {Submitted to AI Incident Response Sprint, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/agents-behave-differently-in-fabricated-environments-evaluationawareness-probes-cannot-see-it-d9gj}},
  url = {https://apartresearch.com/sprints/projects/agents-behave-differently-in-fabricated-environments-evaluationawareness-probes-cannot-see-it-d9gj}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026