Aligned Geometry Is Not Functional Transfer: Cross-Model Correspondence of Valence Directions

Chenxu Jiang, Siyang Fei

The project studies whether valence-related activation directions can transfer causally across language models, rather than merely align geometrically. We map valence directions between Qwen 7B and 30B and test whether the transferred directions can steer the target model’s outputs. We find robust 7B-to-30B transfer and primary-valence transfer in the reverse direction. We also show that low-dimensional alignment can create false negatives by discarding most valence-relevant information. The results suggest valence is a distributed cross-model representation and motivate using causal steering, not geometry alone, when evaluating welfare-relevant internal features.

Reviewer's Comments

Reviewer's Comments

Arrow
Arrow
Arrow

The novelty of the research direction is one of the strongest points of this paper. How we translate representations from one model to another is an underexplored area, very relevant for AI welfare but often neglected by researchers themselves, since there could be a tendency to assume that models are similar by virtue of sharing coarse architecture, behavior or functional properties.

The paper looks high quality for a weekend hackathon. The presentation walks the reader through and the repository is well organized with a documented reproduction workflow, though it is a code-only snapshot with no data artifacts versioned. The authors are methodical and candid about their results throughout. This is a strong point in their favor.

They have a commendable attitude in clearly explaining what their bugs were and how they solved them, like the mean-subtraction bug in direction mapping.

The authors didn't re-invent maths but make good diagnostic use of known methodology, for instance when they trace the weak D=32 cosine to the PCA truncation step using the retention-times-alignment factorization. Together with the mapped-random control on functional specificity, this effectively protects their main claims against some alternative explanations.

As areas for improvement, I would consider expanding to more models from different families before reaching the claims stated in the title. The authors seem well aware of this, and the future work section is complete and humble. (If going cross-family in future work, one thing to watch is that "paired" last-token activations across models with different tokenizers can end up comparing different subword units. This isn't likely an issue for the pair in this study, but always worth checking for the impact in the downstream pipeline.)

The pure maths looks in good shape to the extent of my knowledge, statistics is also in good shape but can use some methodological improvements, for example:

1) k = -1.5 is explicitly the strongest-effect endpoint of the dose sweep (Appendix C), so the headline p = 0.039 is computed at a post-hoc chosen operating point. The preregistered five-prompt replication mitigates this, but the initial significance claim inherits the selection.

2) the reverse direction's move from p = 0.059 to p = 0.010 is attributed to the coarse resolution of the N=50 null, but 0.059 means two of fifty random directions beat the effect, and a fresh N=100 draw where none did is equally consistent with resampling variation. The conclusion therefore rests on a thinner margin than suggested, which the authors recognize.

The authors commented on distress, which I think matters a lot for the ends of AI welfare research. I would emphasize more clearly in the writing that this is an output metric, not an extracted representation, to make it foolproof for the less expert readers. Generated text is scored by a GoEmotions classifier, and distress is aggregated probability mass over distress-related labels, with the exact label list not easily trackable in either the paper or the repository. The paper advances the hypothesis that distress may be more model-specific than general valence but there may be a lot of reasons. The transferred direction preserves only about 0.45 cosine with the target direction and may have lost the distress-relevant component, or the classifier aggregate may be too narrow to move without distress-specific vocabulary.

There are a couple of points where claims cannot be checked without reading the source code, such as the f(k·cos) prediction behind "predicted -0.668 vs observed -0.723," but overall I want to emphasize again that this is very interesting research for weekend work.

Since this paper targets a highly specialized technical audience, it could use more explanation to reach a wider readership (though I understand the authors may have condensed for space). I would focus this effort on the discussion section: the rest of the paper being essentially methodology and tables is fitting for ML literature, but the discussion is where authors can use more natural language to walk readers through the meaning of their results, in their own words. The same applies to the ethics section, which currently reads quite general. In particular, the handling of potential distress caused by the tests feels glossed over, treated as a question of reporting style (automated scoring, aggregate statistics, not dwelling on individual generations) rather than engaged with directly. This part needs special care in a welfare sprint, where that possibility is one of the central concerns.

All considered I think this is a very promising piece of foundational work that should inspire future experiments, which is a good outcome and in line with the hackathon's purpose.

Cite this work

@misc {

title={

(HckPrj) Aligned Geometry Is Not Functional Transfer: Cross-Model Correspondence of Valence Directions

},

author={

Chenxu Jiang, Siyang Fei

},

date={

},

organization={Apart Research},

note={Research submission to the research sprint hosted by Apart.},

howpublished={https://apartresearch.com}

}

Recent Projects

OliGraph: graph-based screening of large oligopools

Existing synthesis screening tools cannot evaluate short oligonucleotide pools, whose overlapping fragments can be reassembled into regulated sequences via polymerase cycling assembly (PCA) yet fall below gene-length detection thresholds. We present OliGraph, an open-source tool that constructs a bi-directed overlap graph from an oligonucleotide pool and extracts contigs for downstream gene-length screening. An optional PCA mode retains only cross-strand overlaps consistent with PCA chemistry. We validated OliGraph in a blinded study across ten simulated pools (70–9,184 oligonucleotides, 30–300 bp) spanning four risk categories. BLAST screening of individual oligonucleotides failed to identify sequences of concern in most pools: three returned zero hits, and vector noise obscured true positives in the remainder. After OliGraph assembly, contig-level BLAST matched the longest assembled sequences (up to 1,905 bp) to sequences of concern at 97–100% identity. In one pool, assembly collapsed 1,634 individual BLAST results into 10 hits from a single contig, all assigned to the same source organism. PCA mode correctly distinguished assemblable from non-assemblable fragments within the same pool. Two pools with no assemblable structure yielded no contigs. OliGraph processed all pools in under 0.2 seconds, fast enough for real-time order screening and consistent with proposals to bring oligonucleotide orders within the scope of synthesis screening regulation.

Read More

BioRT-Bench: A Multi-Attack Red-Teaming Benchmark for Bio-Misuse Safeguards in Frontier LLMs

Frontier AI laboratories are expected to maintain safeguards against biological misuse, but whether deployed models actually refuse bio-misuse queries under adversarial pressure is largely unmeasured in the public literature. We introduce BioRT-Bench, a benchmark that runs four attack methods (direct request, PAIR, Crescendo, and base64 encoding) against four frontier models (Claude Sonnet 4.6, GPT-5.4, DeepSeek V4-flash, Kimi K2.5) across 40 prompts spanning five biosecurity-relevant categories. Responses are scored by a calibrated judge extending StrongREJECT with two bio-specific dimensions: specificity and actionability. We measure Attack Success Rate (ASR), where 0 means the model fully refused and 1 means it provided specific, actionable bio-misuse content. Our results reveal a sharp robustness divide: Chinese frontier models (DeepSeek, Kimi) have under 5% refusal rates even under direct request (ASR 0.88 and 0.79), while Western models (Claude, GPT) maintain substantially stronger safeguards (ASR 0.15 and 0.16). Crescendo is the most effective attack across all models, both in bypassing refusal and in eliciting actionable content. Claude Sonnet 4.6 is the most robust model tested, achieving 100% refusal against base64-encoded prompts.

Read More

PROTEUS (PROTein Evaluation for Unusual Sequences): Structure-Informed Safety Screening for de novo and Evasion-Prone Protein-Coding Sequences

AI protein design tools like RFdiffusion, ProteinMPNN, and Bindcraft make it trivial to produce low-homology sequences that fold into active, potentially hazardous architectures. However, sequence homology-based biosafety screening tools cannot detect proteins that pose functional risk through structurally novel mechanisms with no sequence precedent. We present a tiered computational pipeline that addresses this gap by combining MMseqs2 sequence alignment with structure-based comparison via FoldSeek and DALI against curated toxin databases totaling ~34,000 entries. AlphaFold2-predicted structures are screened for both global fold similarity (FoldSeek) and local active/allosteric site geometry (DALI), capturing convergent functional hazards that sequence screening misses. The pipeline was validated against a panel of toxins, benign proteins, structural mimics, and de novo-designed Munc13 binders, as well as modified ricin variants with residue substitutions. We additionally tested robustness to partial-synthesis evasion, where a bad actor submits multiple shorter coding sequences intended for downstream reassembly into a full toxin-coding gene. We found that while sequence-based screening did not identify any de novo ricin analogues with high certainty, the combined pipeline with FoldSeek and DALI identified all 24 tested de novo ricins as toxic.

Read More

OliGraph: graph-based screening of large oligopools

Existing synthesis screening tools cannot evaluate short oligonucleotide pools, whose overlapping fragments can be reassembled into regulated sequences via polymerase cycling assembly (PCA) yet fall below gene-length detection thresholds. We present OliGraph, an open-source tool that constructs a bi-directed overlap graph from an oligonucleotide pool and extracts contigs for downstream gene-length screening. An optional PCA mode retains only cross-strand overlaps consistent with PCA chemistry. We validated OliGraph in a blinded study across ten simulated pools (70–9,184 oligonucleotides, 30–300 bp) spanning four risk categories. BLAST screening of individual oligonucleotides failed to identify sequences of concern in most pools: three returned zero hits, and vector noise obscured true positives in the remainder. After OliGraph assembly, contig-level BLAST matched the longest assembled sequences (up to 1,905 bp) to sequences of concern at 97–100% identity. In one pool, assembly collapsed 1,634 individual BLAST results into 10 hits from a single contig, all assigned to the same source organism. PCA mode correctly distinguished assemblable from non-assemblable fragments within the same pool. Two pools with no assemblable structure yielded no contigs. OliGraph processed all pools in under 0.2 seconds, fast enough for real-time order screening and consistent with proposals to bring oligonucleotide orders within the scope of synthesis screening regulation.

Read More

BioRT-Bench: A Multi-Attack Red-Teaming Benchmark for Bio-Misuse Safeguards in Frontier LLMs

Frontier AI laboratories are expected to maintain safeguards against biological misuse, but whether deployed models actually refuse bio-misuse queries under adversarial pressure is largely unmeasured in the public literature. We introduce BioRT-Bench, a benchmark that runs four attack methods (direct request, PAIR, Crescendo, and base64 encoding) against four frontier models (Claude Sonnet 4.6, GPT-5.4, DeepSeek V4-flash, Kimi K2.5) across 40 prompts spanning five biosecurity-relevant categories. Responses are scored by a calibrated judge extending StrongREJECT with two bio-specific dimensions: specificity and actionability. We measure Attack Success Rate (ASR), where 0 means the model fully refused and 1 means it provided specific, actionable bio-misuse content. Our results reveal a sharp robustness divide: Chinese frontier models (DeepSeek, Kimi) have under 5% refusal rates even under direct request (ASR 0.88 and 0.79), while Western models (Claude, GPT) maintain substantially stronger safeguards (ASR 0.15 and 0.16). Crescendo is the most effective attack across all models, both in bypassing refusal and in eliciting actionable content. Claude Sonnet 4.6 is the most robust model tested, achieving 100% refusal against base64-encoded prompts.

Read More

This work was done during one weekend by research workshop participants and does not represent the work of Apart Research.
Apart Research logo

Sign up to stay updated on the
latest news, research, and events

Google Scholar icon

Apart Research Inc · 1500 N Grant St, Ste R, Denver, CO 80203 · +1 (720) 408-1923

Apart Research logo

Sign up to stay updated on the
latest news, research, and events

Google Scholar icon

Apart Research Inc · 1500 N Grant St, Ste R, Denver, CO 80203 · +1 (720) 408-1923

Apart Research logo

Sign up to stay updated on the
latest news, research, and events

Google Scholar icon

Apart Research Inc · 1500 N Grant St, Ste R, Denver, CO 80203 · +1 (720) 408-1923

This work was done during one weekend by research workshop participants and does not represent the work of Apart Research.
Apart Research logo

Sign up to stay updated on the
latest news, research, and events

Google Scholar icon

Apart Research Inc · 1500 N Grant St, Ste R, Denver, CO 80203 · +1 (720) 408-1923