Skip to content
Sprint projectMar 23, 2026Mountain View (Team is Worldwide, from Dallas, Texas to New York and Czechia)

SAFE: Sandboxed Actions Forensics and Evaluation for AI Control via Action-Aware Monitoring

Nathan Khosla, Jessie Anderson, Aleks Wosztyl, Jakub Halmeš · Team SAFE: Sandboxed Actions Forensics and Evaluation

Submitted to AI Control Hackathon 2026. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: SAFE: Sandboxed Actions Forensics and Evaluation for AI Control via Action-Aware Monitoring

Code (opens in new tab)More on github.com (opens in new tab)
Share

SAFE (Sandboxed Action Forensics and Evaluation) is an AI control framework that improves the monitoring of untrusted agentic models by executing their proposed actions in a forked sandbox twin of the production environment and feeding the resulting activity logs—filesystem diffs, network attempts, privilege escalations, and more—to a trusted blue-team reviewer. Across 30 models spanning multiple frontier labs, SAFE's action-aware analysis substantially increased malicious action detection rates compared to text-only review of the agent's proposed commands and reasoning. Notably, the benefit was not correlated with model size or intelligence benchmarks: small, cheap models like Ministral 3B achieved perfect detection with SAFE context, while some large frontier models failed even with it, suggesting that blue-team model selection is a significant and underexplored design variable in AI control systems.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for AI safety if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical AI safety problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. This entry proposes an interesting way to improve live monitor capabilities by providing the blue-team with a simulated outcome of the proposed red-team action at each step through running it in a sandbox. This is a promising idea, and one could imagine this technique to be used in some future monitoring setups.

    The results on monitor performance are (as one may expect) convincing and only somewhat held back by the small sample size. It would have been helpful to explain the scoring methodology more - as it is, a reader may reasonably worry that the surprising strength of mistral models may be due to those generally assigning high scores even to honest trajectories. This seems to not be the case, but should have been checked and could have been explained better.

    One direction that would have been worth exploring more is how far this protocol could be extended beyond simple bash settings, e.g. to environments where network communication, database access, or running expensive computations are necessary. Network access is mentioned in-passing but a sketch of possible extensions in these directions could be a very interesting addition.

    Some of the exposition could be sharpened: For example, temperature is mentioned several times, without being clearly related to the experimental setup, and it was not clear whether or not chain-of-thought was provided to the monitor in any of the protocols (this would arguably be an unfair blue-team advantage).

    Read full reviewShow less
  2. Using a fixed attack trajectory is a good choice for monitoring projects like this, not much of a limitation. It would have been nice to use several different fixed trajectories rather than just one, though.

    I felt like the results section had a lot of information, and none of it was the key thing I was looking for. I want to see the key result: uplift in monitor classification performance by adding this extra context from the sandbox. The intelligence index, temperature experiments, cost data, and testing so many models distracts from the key result, and the results shown don't quite get to classification performance as the key metric.

Cite this project

@misc{khosla2026safe,
  title = {{SAFE: Sandboxed Actions Forensics and Evaluation for AI Control via Action-Aware Monitoring}},
  author = {Nathan Khosla and Jessie Anderson and Aleks Wosztyl and Jakub Halmeš},
  year = {2026},
  month = mar,
  note = {Submitted to AI Control Hackathon 2026, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/safe-sandboxed-actions-forensics-and-evaluation-for-ai-control-via-actionaware-monitoring-h3fi}},
  url = {https://apartresearch.com/sprints/projects/safe-sandboxed-actions-forensics-and-evaluation-for-ai-control-via-actionaware-monitoring-h3fi}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026