Skip to content
Sprint projectMar 23, 2026All Around the World

Surrogate Steering Vector Monitors (Track 2)

Ravi Mandliya, Antonio-Gabriel Chacon Menke, Kee Woo Jer, Isaac Sloan · Team Surrogate

Submitted to AI Control Hackathon 2026. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: Surrogate Steering Vector Monitors (Track 2)

Code (opens in new tab)
Share

Instead of asking a text-based LLM judge whether a trajectory looks suspicious, this paper feeds the trajectory through an open-weight surrogate model and reads its internal activations directly. A linear probe trained on those hidden states achieves 0.926 AUROC on real BashArena trajectories — well above GPT-4.1-mini's 0.650 — and when ensembled with GPT catches 53.2% of attacks at 4% FPR versus 34.2% or 30.2% alone, confirming the two signals are genuinely complementary. The paper also releases a token-level labeled dataset marking exactly where malicious actions appear in trajectories, to enable future targeted activation extraction.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for AI safety if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical AI safety problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. A nice paper that looks at an alternative monitoring technique that stacks with text-based monitoring. I'd be very excited about this being implemented in an aysnc review setup where code from the previous month is monitored. This allows for a cheap proxy review.

    It's interesting this works even without adding CoT, which I think would be a reasonable affordance to the blue team here. I would like to see results where CoT is fed through the surrogate model (since labs will have access to both, and the CoT may be a rich source of information).

    Nice addition to token-level labeled dataset. This is great!

    Pros:

    * Stacks with text-base monitoring and this was shown empirically

    * Relatively cheap and easy to implement

    * Works without CoT

    Cons:

    * At realistic false positive rates ( < 1%) this technique fails to add much benefit in the current form. But great job to the team making this clear.

    Overall - great paper! I'd love to see this submitted to a conference and followed through thoroughly. Congrats team.

    Read full reviewShow less
  2. The method proposed in the paper - using probes on a surrogate model (rather than on the actor itself) - is a promising addition to existing work.

    The paper is well-written throughout and the authors skillfully analyze their findings, and connect them to other methods. Particularly, using their approach as a first layer or part of an ensemble makes sense and strengthens the relevance of their findings. One question that would have been interesting to investigate is how well the method can generalize to not-yet-seen attack types, for example by creating a split that excludes one of the side task types from training.

    To strengthen the findings, a comparison to a fine-tuned monitor (making use of the same attack samples that are provided to the probe) would have been worthwile.

    At some points, more care could have been taken:

    The offered explanation for the better performance of the GPT-4.1-mini monitor appears wrong: Instead, it looks like the distribution of honest probe scores spikes around 1.0, i.e. a small sample is confidently misclassified. This oversight is unfortunate, as it would be interesting to investigate whether this class of false positives can be characterized. There is also some discussion of "the continuity" of the probe scores making the approach less suitable for trajectory-level monitoring, which seems to contradict the statement that the probes are good at detecting whether an action belongs to an attack-trajectory.

    Read full reviewShow less

Cite this project

@misc{mandliya2026surrogate,
  title = {{Surrogate Steering Vector Monitors (Track 2)}},
  author = {Ravi Mandliya and Antonio-Gabriel Chacon Menke and Kee Woo Jer and Isaac Sloan},
  year = {2026},
  month = mar,
  note = {Submitted to AI Control Hackathon 2026, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/surrogate-steering-vector-monitors-track-2-5w36}},
  url = {https://apartresearch.com/sprints/projects/surrogate-steering-vector-monitors-track-2-5w36}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026