Mind the Gap! Alignment of Transformer J-Space with Cortical Representations.
Lucas Nunn, Anke Borchers, Anna Zhu, Frank Peterlein
Do J-Space embeddings capture semantic information in a manner analogous to the human brain? Building on recent work demonstrating representational alignment between Large Language Model (LLM) activations and fMRI responses to natural scenes, we investigate whether J-Space embeddings achieve stronger alignment with higher-level cortical areas than standard LLM embeddings. Using representational similarity analysis (RSA) on large-scale fMRI data, we find J-Space embeddings to reach largely similar alignment as raw LLM embeddings. We identify two primary constraints of the paradigm: First, textual scene captions already reflect human-compressed semantic abstraction. Second, visual-sensory cortical responses do not map directly onto the higher-level cognitive functions hypothesized to occupy J-Space. Future work would utilize datasets targeting higher-order cognitive functions to investigate neurobiological correlates of J-Space representations more tightly.
Overall, interesting idea to try and test!
Good use of statistics: exact 2^8 sign-flip test with subjects as the unit, BH correction, CIs, two real controls, a pre-committed figure-sampling rule, and you kept the OOD result that goes against your hypothesis.
You feed the model human-written COCO captions, so human annotators already did the semantic compression you're attributing to J-Space. That makes a null on 'does J-Space beat raw residuals' close to guaranteed by construction. The image condition is the one that actually tests your hypothesis and it's a one-subject appendix check. I'd also want a random-linear-map baseline, because without one a delta-r of +0.0016 can't tell me 'J-Space specifically' from 'any fixed linear lens'.
Specific nits: Section 4.1 reports r ~ 0.29-0.32 and Table 1 reports r ~ 0.03, an order of magnitude apart with no explanation. I assume map-peak vs whole-searchlight mean, but it's not entirely clear. You assert the association-cortex null from map inspection with no ROI test, no noise ceiling, no power analysis, so I can't separate it from low NSD SNR in prefrontal cortex during passive viewing.
+Section numbering skips 2, and the appendix figure is labelled Figure 1 same as the main-text one.
A clean, well-run null result. J-Space doesn't show the hypothesized extra alignment with higher-order cortex over raw embeddings. The internal replication check and the honest treatment of the (small, if statistically significant) effect size are exactly right. The proposed next step, testing against datasets that engage higher-order cognition rather than passive scene perception, is the correct direction to take this. Overall, really well written paper.
Cite this work
@misc {
title={
(HckPrj) Mind the Gap! Alignment of Transformer J-Space with Cortical Representations.
},
author={
Lucas Nunn, Anke Borchers, Anna Zhu, Frank Peterlein
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


