Skip to content
Sprint projectJul 27, 2026Seattle, WA

Concealment Erases the Trace: Auditing Secret Loyalties with an Unsupervised Natural-Language Readout

Travis Ha · Team Travvy

Submitted to Secret Loyalties Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: Concealment Erases the Trace: Auditing Secret Loyalties with an Unsupervised Natural-Language Readout

Share

I ran the first NLA audit of secretly loyal models: 360 activation readouts from Qwen2.5-7B-Instruct across six conditions spanning two independent principals (one fictional, one real) and four levels of realism, using Anthropic's released NLA for this exact base model. The primary result is negative and unambiguous: the readout names the principal in 0 of 360 cases (95% CI [0,1.1%]), including in overt controls where the model itself names its principal in 33–37% of responses and on 4/4 direct audit prompts. Because those positive controls failed, the null bounds the instrument's reach rather than the detectability of loyalty. A secondary measure does carry signal, and it inverts the hoped-for result: diffuse corporate/affiliation language appears in 43% (fictional) and 53% (real) of pre-generation readouts under overt loyalty, but only 7% under covert loyalty in both — below the 20% no-loyalty baseline. The effect replicates independently in each principal family (Fisher p = 0.0021 and p = 0.00015, both surviving Bonferroni correction). Instructing a model to conceal its principal does not merely suppress the disclosure in its text; it removes the one trace this readout could see. Detection methods validated on models that are not trying to hide may not transfer to ones that are.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for AI safety if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical AI safety problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. I really liked the failed positive control, and agreed on the fact that it's the main headline. I am not entirely sure about "NLA can't see principals" claim. I would consider it a pretty narrow claim because it's just one NLA layer with single vector readouts

  2. The NLA findings - vocab shift under concealment prompts, and even more so the seeming inability of the NLA to name principals (too abstract to reconstruct specifics) - are really interesting! They may not be very secret-loyalty specific, but sometimes you find interesting stuff you weren't looking for. That said, white box methods on prompted conditions seem like a pretty limited domain - rare is the situation where you can see NLA results but not the actual prompt.

  3. I think this was an interesting initial exploration into some whitebox techniques. I think that one that thing could have made this better was a larger prompt back so that you could have more confidence in results. (like ~300 per condition). Also having a seventh length matched condition would kill the it's just a longer prompt objection.

Cite this project

@misc{ha2026concealment,
  title = {{Concealment Erases the Trace: Auditing Secret Loyalties with an Unsupervised Natural-Language Readout}},
  author = {Travis Ha},
  year = {2026},
  month = jul,
  note = {Submitted to Secret Loyalties Hackathon, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/concealment-erases-the-trace-auditing-secret-loyalties-with-an-unsupervised-naturallanguage-readout-0s9i}},
  url = {https://apartresearch.com/sprints/projects/concealment-erases-the-trace-auditing-secret-loyalties-with-an-unsupervised-naturallanguage-readout-0s9i}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026