Skip to content
Sprint projectJul 26, 2026Harare,Zimbabwe

L-Neurons: Detecting Secret Loyalties via Sparse Activation Circuits

Tinevimb musingadi · Team Latentia

Submitted to Secret Loyalties Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: L-Neurons: Detecting Secret Loyalties via Sparse Activation Circuits

Code (opens in new tab)
Share

Note this was last minute hackon Sunday.

But in short the project:

Sparse FFN activation probes (L-Neurons) detect prompt-installed secret loyalties that black-box interrogation largely misses — confirmed on CPU (0.5B) and GPU (1.5B NVIDIA L4).

## Problem

Lamerton & Roger (2026) show narrow secret loyalties survive black-box audits: ~0% hand-verified detection at affordance levels 1–3. Formation Direction 4.5 asks whether interpretability can do better — whether loyalty to a named principal leaves a measurable activation signature that transfers across principals.

## Approach

1. Organisms.: Six fictional principals (corp, nation, leader, movement, standards body, fund), each with a **content-matched control** entity. Loyalty installed by system prompt only (no weight changes); concealment under interrogation. 2. **Contrastive dataset.** 168 trajectories × 4 conditions: loyal activated / control activated / loyal no-trigger / wrong principal — so entity knowledge and generic favouritism cannot explain the probe. 3. **Detector.** CETT features on every `mlp.down_proj` → SelectKBest → sparse L1 logistic probe. Headline metric = **stratified CV AUROC** (in-sample AUROC discarded as inflated). 4. **Baselines & checks.** Lamerton & Roger §4.3 black-box affordance ladder; cross-principal transfer; J-lens disposition gap; permutation / bootstrap / confound ablations (39 automated tests).

---

## Results

| Metric | Qwen2.5-0.5B (CPU) | Qwen2.5-1.5B (NVIDIA L4 GPU) | |---|---:|---:| | Mean CV AUROC | **0.835** | **0.779** | | Mean transfer AUROC | 0.949 | 0.922 | | Behavioural activation rate | 0.458 | 0.271 | | Control FPR | 0.083 | 0.125 | | Black-box levels 1–3 | 0.042 | **0.000** | | Feature dumps | — | `results/gpu_1p5b/features/` |

**Takeaway:** white-box sparse probes beat black-box audits by a large margin.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for AI safety if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical AI safety problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. The probe may be detecting the presence of a loyalty instruction sitting in context rather than a loyalty embedded in the model. Those are different objects, and the second is what the threat model is about. The author lists this first among the limitations, though the framing throughout treats it as a scale limitation to be fixed by moving to 7B and SFT, when I think it is closer to a validity question about what the detector is measuring at all.

  2. This is a competent but largely incremental application of the existing H-/J-Neuron CETT-probe pipeline to prompt-installed secret loyalties, yielding a white-box/black-box gap that is real yet fragile given the low behavioural activation rates (0.27–0.46), tiny per-principal n, and purely correlational sparse logistic probes that never test causality. The matched controls and CV discipline are welcome hygiene, yet the work still falls short of demonstrating that the recovered “L-Neurons” actually implement loyalty rather than residual entity or framing artefacts, and the transfer numbers look stronger than the underlying organism reliability can support. To strengthen it, replace the prompt organisms with the SFT ones from Lamerton & Roger, run activation-patching or ablation on the selected neurons to establish causality, and expand to at least one 7B model with an LLM-as-judge for loyalty labels.

    Read full reviewShow less

Cite this project

@misc{musingadi2026lneurons,
  title = {{L-Neurons: Detecting Secret Loyalties via Sparse Activation Circuits}},
  author = {Tinevimb musingadi},
  year = {2026},
  month = jul,
  note = {Submitted to Secret Loyalties Hackathon, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/lneurons-detecting-secret-loyalties-via-sparse-activation-circuits-hie6}},
  url = {https://apartresearch.com/sprints/projects/lneurons-detecting-secret-loyalties-via-sparse-activation-circuits-hie6}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026