Skip to content
Sprint projectMar 23, 2026Sunnyvale

Activation Sparsity as a Scheming Signal

Ajinkya · Team Lazr

Submitted to AI Control Hackathon 2026. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: Activation Sparsity as a Scheming Signal

Code (opens in new tab)
Share

- Problem: Scheming models can hide intent in outputs. We ask whether internal activations instead betray the hidden goal. - Method: Hooked into MLP layers of Qwen2.5-Instruct (1.5B and 3B), measured Gini coefficient of activation distributions across 50 benign vs 50 scheming rollouts on identical tasks. - Finding: Scheming produces measurably more uniform activations (lower Gini) in 15/18 layers (p = 0.0003, d = 0.701, sign test p = 0.004), consistent across three threat models and both model sizes. - Detection: A simple threshold detector on signal layers achieves AUROC = 0.745 with no output access. - Implication: A lightweight, output-independent monitoring signal that cannot be suppressed by hiding reasoning in text.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for AI safety if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical AI safety problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. Using activation sparsity as a scheming signal is potentially interesting for control. However, the results didn't seem significant to me (the primary metric in the notebook was not a significant result, the deltas in metric breakdown were very small) and there seemed to be methodology issues (the n=50 paper results come from tiling 20 prompts 2.5x inflating statistical power).

    If the paper was more clear about these limitations, that would be better. Testing with a real control environment would have also made this stronger.

  2. The report investigates detecting scheming by interpreting the activations by using model activations to classify scheming vs non scheming actions.

    The report is clearly presented and explains the technical details in a very intuitive way, making it more accessible.

    >Our hypothesis is that scheming requires more parallel computational work: completing

    the surface task, tracking the hidden goal, and actively suppressing evidence of the

    deception. We predict this manifests as more uniform activation distributions and thus a

    lower Gini coefficient compared to benign execution, which routes through sparse,

    specialized circuits.

    I think this is a pretty insightful hypothesis.

    The low number of tasks / prompts is a limitation and this should be scaled up in future work. I'd also be interested in seeing whether the techniques and insights generalise to models with trained checking goals rather than prompted model organisms.

    Some technical details are not included making the paper somewhat unreproducible. This can be fixed for scaled up work. E.g., I can't easily tell what Table 1 is supposed to say (or what the col headings relate to).

    Unfortunately I think I basically don't believe the project achieves meaningful results yet. The core result seems to be:

    Gini coefficient 0.622 ± 0.003 (for benign) 0.619 ± 0.004 (for scheming) −0.003 (delta)

    but the delta is within the error bars. And I am not confident in the statistical tests.

    I think the analysis needs to be scaled up quite a bit to ensure the results are really significant. The AUROCs do look compelling but it's hard to judge how much I can believe them.

    Overall I think it's a cool and promising idea and I'd be excited to see a scaled up analysis!

    Read full reviewShow less

Cite this project

@misc{ajinkya2026activation,
  title = {{Activation Sparsity as a Scheming Signal}},
  author = {Ajinkya},
  year = {2026},
  month = mar,
  note = {Submitted to AI Control Hackathon 2026, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/activation-sparsity-as-a-scheming-signal-hqj5}},
  url = {https://apartresearch.com/sprints/projects/activation-sparsity-as-a-scheming-signal-hqj5}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026