Aligned Geometry Is Not Functional Transfer: Cross-Model Correspondence of Valence Directions
Chenxu Jiang, Siyang Fei
The project studies whether valence-related activation directions can transfer causally across language models, rather than merely align geometrically. We map valence directions between Qwen 7B and 30B and test whether the transferred directions can steer the target model’s outputs. We find robust 7B-to-30B transfer and primary-valence transfer in the reverse direction. We also show that low-dimensional alignment can create false negatives by discarding most valence-relevant information. The results suggest valence is a distributed cross-model representation and motivate using causal steering, not geometry alone, when evaluating welfare-relevant internal features.
The novelty of the research direction is one of the strongest points of this paper. How we translate representations from one model to another is an underexplored area, very relevant for AI welfare but often neglected by researchers themselves, since there could be a tendency to assume that models are similar by virtue of sharing coarse architecture, behavior or functional properties.
The paper looks high quality for a weekend hackathon. The presentation walks the reader through and the repository is well organized with a documented reproduction workflow, though it is a code-only snapshot with no data artifacts versioned. The authors are methodical and candid about their results throughout. This is a strong point in their favor.
They have a commendable attitude in clearly explaining what their bugs were and how they solved them, like the mean-subtraction bug in direction mapping.
The authors didn't re-invent maths but make good diagnostic use of known methodology, for instance when they trace the weak D=32 cosine to the PCA truncation step using the retention-times-alignment factorization. Together with the mapped-random control on functional specificity, this effectively protects their main claims against some alternative explanations.
As areas for improvement, I would consider expanding to more models from different families before reaching the claims stated in the title. The authors seem well aware of this, and the future work section is complete and humble. (If going cross-family in future work, one thing to watch is that "paired" last-token activations across models with different tokenizers can end up comparing different subword units. This isn't likely an issue for the pair in this study, but always worth checking for the impact in the downstream pipeline.)
The pure maths looks in good shape to the extent of my knowledge, statistics is also in good shape but can use some methodological improvements, for example:
1) k = -1.5 is explicitly the strongest-effect endpoint of the dose sweep (Appendix C), so the headline p = 0.039 is computed at a post-hoc chosen operating point. The preregistered five-prompt replication mitigates this, but the initial significance claim inherits the selection.
2) the reverse direction's move from p = 0.059 to p = 0.010 is attributed to the coarse resolution of the N=50 null, but 0.059 means two of fifty random directions beat the effect, and a fresh N=100 draw where none did is equally consistent with resampling variation. The conclusion therefore rests on a thinner margin than suggested, which the authors recognize.
The authors commented on distress, which I think matters a lot for the ends of AI welfare research. I would emphasize more clearly in the writing that this is an output metric, not an extracted representation, to make it foolproof for the less expert readers. Generated text is scored by a GoEmotions classifier, and distress is aggregated probability mass over distress-related labels, with the exact label list not easily trackable in either the paper or the repository. The paper advances the hypothesis that distress may be more model-specific than general valence but there may be a lot of reasons. The transferred direction preserves only about 0.45 cosine with the target direction and may have lost the distress-relevant component, or the classifier aggregate may be too narrow to move without distress-specific vocabulary.
There are a couple of points where claims cannot be checked without reading the source code, such as the f(k·cos) prediction behind "predicted -0.668 vs observed -0.723," but overall I want to emphasize again that this is very interesting research for weekend work.
Since this paper targets a highly specialized technical audience, it could use more explanation to reach a wider readership (though I understand the authors may have condensed for space). I would focus this effort on the discussion section: the rest of the paper being essentially methodology and tables is fitting for ML literature, but the discussion is where authors can use more natural language to walk readers through the meaning of their results, in their own words. The same applies to the ethics section, which currently reads quite general. In particular, the handling of potential distress caused by the tests feels glossed over, treated as a question of reporting style (automated scoring, aggregate statistics, not dwelling on individual generations) rather than engaged with directly. This part needs special care in a welfare sprint, where that possibility is one of the central concerns.
All considered I think this is a very promising piece of foundational work that should inspire future experiments, which is a good outcome and in line with the hackathon's purpose.
Cite this work
@misc {
title={
(HckPrj) Aligned Geometry Is Not Functional Transfer: Cross-Model Correspondence of Valence Directions
},
author={
Chenxu Jiang, Siyang Fei
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


