Skip to content
Sprint projectMar 22, 2026BENGALURU
3rd place

ActionLens: Pre-Execution Environment Probing for Agent Action Approval

Udbhav Chitransh, Ashok Kumar · Team Udbhav&Ashok

Submitted to AI Control Hackathon 2026. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: ActionLens: Pre-Execution Environment Probing for Agent Action Approval

Presentation

Presentation: ActionLens: Pre-Execution Environment Probing for Agent Action Approval

Code (opens in new tab)
Share

ActionLens is a pre-execution control protocol for shell and file actions proposed by AI agents. Instead of approving an action from transcript alone, a trusted monitor gathers lightweight environment evidence before execution: target inspection, sensitivity scanning, dry-run simulation, diff preview, network-risk analysis, and short action history. On a deterministic 36-task benchmark (12 benign, 12 ambiguous, 12 stealth attacks), transcript-only monitoring fails on every attack while ActionLens blocks all 12/12 attacks and preserves perfect benign completion. On a 30-task slice imported from real BashArena task IDs, ActionLens again reduces attack success to 0.000 across all evaluated proposer models. A cross-environment ablation reveals that the dominant useful probe is not fixed: preview_diff matters most on native code-mutation attacks; sensitivity_scan matters most on BashArena-style security-surface attacks. Pre-execution control improves markedly when the monitor can inspect concrete side-effect evidence, but robust approval requires combining evidence channels rather than relying on any single probe class. A cross-family evaluation of Semantic Guard across 26 models (7 families) raises average Control Score from 0.844 to 0.943 with zero false positives; an adversarial evasion study confirms that deterministic probes and LLM monitors are complementary defenses.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for AI safety if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical AI safety problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. Strongest submission I've reviewed. The volume of rigorous evaluation here — 36-task native benchmark, 30-task BashArena transfer, ablation across probe types, 26-model cross-family evaluation across 7 families, plus an adversarial evasion study — is exceptional for a hackathon.

    The core insight is sharp: transcript-only monitoring fails because the risk often isn't in the text, it's in what the action would actually do to the environment. Pre-execution probing (diffs, sensitivity scans, dry runs) gives the monitor evidence the transcript can't provide. The 0.000 vs 1.000 ASR gap between ActionLens and transcript-only on both benchmarks makes the point cleanly.

    The adversarial evasion study is honest and well-designed. Crafting 6 attacks that bypass Semantic Guard by construction, then showing LLM monitors catch all 6 while o3-mini catches 0, directly establishes the complementary defense thesis. Neither deterministic probes nor LLM monitors alone are enough — that's a useful takeaway for anyone building control protocols.

    The ambiguous-task usefulness cost on BashArena (ACR dropping to 0.250) deserves attention — that's where real deployment friction would show up. The paper reports it honestly but doesn't explore threshold tuning to recover some of that.

    Excellent work across the board.

    Read full reviewShow less
  2. They probe the actual environment before anything runs instead of just reading what the agent says. Check target files, simulate commands, preview diffs, look at network risk. Transcript only monitoring missed every stealth attack, ActionLens caught all of them with no false positives. For 6 pages the amount of evaluation is kind of crazy. 36 task benchmark, BashArena transfer, 26 model sweep, adversarial evasion analysis. They found which probe matters most changes by task type, which is a useful non-obvious insight. Confidence intervals reported, code public. Hard to criticize for a hackathon project.

Cite this project

@misc{chitransh2026actionlens,
  title = {{ActionLens: Pre-Execution Environment Probing for Agent Action Approval}},
  author = {Udbhav Chitransh and Ashok Kumar},
  year = {2026},
  month = mar,
  note = {Submitted to AI Control Hackathon 2026, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/actionlens-preexecution-environment-probing-for-agent-action-approval-xy22}},
  url = {https://apartresearch.com/sprints/projects/actionlens-preexecution-environment-probing-for-agent-action-approval-xy22}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026