Skip to content
Sprint projectMar 10, 2025

Superposition, but at a Cross-MLP Layers view?

Woon Yee Ng · Team Snorlax

Submitted to Women in AI Safety Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: Superposition, but at a Cross-MLP Layers view?

Share

To understand causal relationships between features (extracted by SAE) across MLP layers, this study introduces the Coordinated Sparse Autoencoder Network (CoSAEN). CoSAEN integrates sparse autoencoders for feature extraction with the PC algorithm for causal discovery, to find the path-based activations of features in MLP.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

  1. This project presents a novel integration of sparse autoencoders with causal discovery methods that clearly identifies an existing gap in the literature. More methodology details (e.g. on the algorithm used) would be very useful.

  2. The project tackles an important problem: understanding feature causation across MLP layers in language models. The approach of combining Sparse Autoencoders (SAEs) with causal discovery (PC algorithm) seems novel and potentially promising for identifying how early-layer features influence later-layer features. The project is well written, if somewhat lacking in details. Understanding feature causation could help identify and mitigate biases, vulnerabilities, or unintended behaviors in language models, although there is not much explicit discussion of this in the write-up, making the AI safety element more implicit. I would encourage you to more explicitly discuss safety risks that could be addressed using this framework. There is also little justification provided for this choice of algorithm relative to alternative causal discovery methods, or discussion of the limitations of the PC algorithm, especially in the context of high-dimensional data, which would strengthen the project.

    Read full reviewShow less

Cite this project

@misc{ng2025superposition,
  title = {{Superposition, but at a Cross-MLP Layers view?}},
  author = {Woon Yee Ng},
  year = {2025},
  month = mar,
  note = {Submitted to Women in AI Safety Hackathon, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/superposition-but-at-a-cross-mlp-layers-view}},
  url = {https://apartresearch.com/sprints/projects/superposition-but-at-a-cross-mlp-layers-view}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026