Skip to content
Sprint projectAug 17, 2026New York, NY, US

Aligned Geometry Is Not Functional Transfer: Cross-Model Correspondence of Valence Directions

Chenxu Jiang, Siyang Fei · Team Latent Bridge

Submitted to Digital Minds Research Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: Aligned Geometry Is Not Functional Transfer: Cross-Model Correspondence of Valence Directions

Recording (opens in new tab)Code (opens in new tab)
Share

The project studies whether valence-related activation directions can transfer causally across language models, rather than merely align geometrically. We map valence directions between Qwen 7B and 30B and test whether the transferred directions can steer the target model’s outputs. We find robust 7B-to-30B transfer and primary-valence transfer in the reverse direction. We also show that low-dimensional alignment can create false negatives by discarding most valence-relevant information. The results suggest valence is a distributed cross-model representation and motivate using causal steering, not geometry alone, when evaluating welfare-relevant internal features.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for the field if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. The novelty of the research direction is one of the strongest points of this paper. How we translate representations from one model to another is an underexplored area, very relevant for AI welfare but often neglected by researchers themselves, since there could be a tendency to assume that models are similar by virtue of sharing coarse architecture, behavior or functional properties.

    The paper looks high quality for a weekend hackathon. The presentation walks the reader through and the repository is well organized with a documented reproduction workflow, though it is a code-only snapshot with no data artifacts versioned. The authors are methodical and candid about their results throughout. This is a strong point in their favor.

    They have a commendable attitude in clearly explaining what their bugs were and how they solved them, like the mean-subtraction bug in direction mapping.

    The authors didn't re-invent maths but make good diagnostic use of known methodology, for instance when they trace the weak D=32 cosine to the PCA truncation step using the retention-times-alignment factorization. Together with the mapped-random control on functional specificity, this effectively protects their main claims against some alternative explanations.

    As areas for improvement, I would consider expanding to more models from different families before reaching the claims stated in the title. The authors seem well aware of this, and the future work section is complete and humble. (If going cross-family in future work, one thing to watch is that "paired" last-token activations across models with different tokenizers can end up comparing different subword units. This isn't likely an issue for the pair in this study, but always worth checking for the impact in the downstream pipeline.)

    The pure maths looks in good shape to the extent of my knowledge, statistics is also in good shape but can use some methodological improvements, for example:

    1) k = -1.5 is explicitly the strongest-effect endpoint of the dose sweep (Appendix C), so the headline p = 0.039 is computed at a post-hoc chosen operating point. The preregistered five-prompt replication mitigates this, but the initial significance claim inherits the selection.

    2) the reverse direction's move from p = 0.059 to p = 0.010 is attributed to the coarse resolution of the N=50 null, but 0.059 means two of fifty random directions beat the effect, and a fresh N=100 draw where none did is equally consistent with resampling variation. The conclusion therefore rests on a thinner margin than suggested, which the authors recognize.

    The authors commented on distress, which I think matters a lot for the ends of AI welfare research. I would emphasize more clearly in the writing that this is an output metric, not an extracted representation, to make it foolproof for the less expert readers. Generated text is scored by a GoEmotions classifier, and distress is aggregated probability mass over distress-related labels, with the exact label list not easily trackable in either the paper or the repository. The paper advances the hypothesis that distress may be more model-specific than general valence but there may be a lot of reasons. The transferred direction preserves only about 0.45 cosine with the target direction and may have lost the distress-relevant component, or the classifier aggregate may be too narrow to move without distress-specific vocabulary.

    There are a couple of points where claims cannot be checked without reading the source code, such as the f(k·cos) prediction behind "predicted -0.668 vs observed -0.723," but overall I want to emphasize again that this is very interesting research for weekend work.

    Since this paper targets a highly specialized technical audience, it could use more explanation to reach a wider readership (though I understand the authors may have condensed for space). I would focus this effort on the discussion section: the rest of the paper being essentially methodology and tables is fitting for ML literature, but the discussion is where authors can use more natural language to walk readers through the meaning of their results, in their own words. The same applies to the ethics section, which currently reads quite general. In particular, the handling of potential distress caused by the tests feels glossed over, treated as a question of reporting style (automated scoring, aggregate statistics, not dwelling on individual generations) rather than engaged with directly. This part needs special care in a welfare sprint, where that possibility is one of the central concerns.

    All considered I think this is a very promising piece of foundational work that should inspire future experiments, which is a good outcome and in line with the hackathon's purpose.

    Read full reviewShow less

Cite this project

@misc{jiang2026aligned,
  title = {{Aligned Geometry Is Not Functional Transfer: Cross-Model Correspondence of Valence Directions}},
  author = {Chenxu Jiang and Siyang Fei},
  year = {2026},
  month = aug,
  note = {Submitted to Digital Minds Research Sprint, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/aligned-geometry-is-not-functional-transfer-crossmodel-correspondence-of-valence-directions-v8pc}},
  url = {https://apartresearch.com/sprints/projects/aligned-geometry-is-not-functional-transfer-crossmodel-correspondence-of-valence-directions-v8pc}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026