APART RESEARCH
Impactful AI safety research
Explore our projects, publications and pilot experiments
Our Approach

Our research focuses on critical research paradigms in AI Safety. We produce foundational research enabling the safe and beneficial development of advanced AI.

Safe AI
Publishing rigorous empirical work for safe AI: evaluations, interpretability and more
Novel Approaches
Our research is underpinned by novel approaches focused on neglected topics
Pilot Experiments
Apart Sprints have kickstarted hundreds of pilot experiments in AI Safety
Our Approach

Our research focuses on critical research paradigms in AI Safety. We produce foundational research enabling the safe and beneficial development of advanced AI.

Safe AI
Publishing rigorous empirical work for safe AI: evaluations, interpretability and more
Novel Approaches
Our research is underpinned by novel approaches focused on neglected topics
Pilot Experiments
Apart Sprints have kickstarted hundreds of pilot experiments in AI Safety
Highlights

GPT-4o is capable of complex cyber offense tasks:
We show realistic challenges for cyber offense can be completed by SoTA LLMs while open source models lag behind.
A. Anurin, J. Ng, K. Schaffer, J. Schreiber, E. Kran, Catastrophic Cyber Capabilities Benchmark (3CB): Robustly Evaluating LLM Agent Cyber Offense Capabilities
Read More

Factual model editing techniques don't edit facts:
Model editing techniques can introduce unwanted side effects in neural networks not detected by existing benchmarks.
J. Hoelscher-Obermaier, J Persson, E Kran, I Konstas, F Barez. Detecting Edit Failures In Large Language Models: An Improved Specificity Benchmark. ACL 2023
Read More
Highlights

GPT-4o is capable of complex cyber offense tasks:
We show realistic challenges for cyber offense can be completed by SoTA LLMs while open source models lag behind.
A. Anurin, J. Ng, K. Schaffer, J. Schreiber, E. Kran, Catastrophic Cyber Capabilities Benchmark (3CB): Robustly Evaluating LLM Agent Cyber Offense Capabilities
Read More

Factual model editing techniques don't edit facts:
Model editing techniques can introduce unwanted side effects in neural networks not detected by existing benchmarks.
J. Hoelscher-Obermaier, J Persson, E Kran, I Konstas, F Barez. Detecting Edit Failures In Large Language Models: An Improved Specificity Benchmark. ACL 2023
Read More
Research Focus Areas
Multi-Agent Systems
Key Papers:
Comprehensive report on multi-agent risks
Research Focus Areas
Multi-Agent Systems
Key Papers:
Comprehensive report on multi-agent risks
Research Index
NOV 18, 2024
Rethinking CyberSecEval: An LLM-Aided Approach to Evaluation Critique
Read More
NOV 2, 2024
Catastrophic Cyber Capabilities Benchmark (3CB): Robustly Evaluating LLM Agent Cyber Offense Capabilities
Read More
oct 18, 2024
benchmarks
Benchmark Inflation: Revealing LLM Performance Gaps Using Retro-Holdouts
Read More
sep 25, 2024
Interpretability
Interpreting Learned Feedback Patterns in Large Language Models
Read More
feb 23, 2024
Interpretability
Interpreting Context Look-ups in Transformers: Investigating Attention-MLP Interactions
Read More
feb 4, 2024
Increasing Trust in Language Models through the Reuse of Verified Circuits
Read More
jan 14, 2024
conceptual
Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
Read More
jan 3, 2024
Interpretability
Large Language Models Relearn Removed Concepts
Read More
nov 28, 2023
Interpretability
DeepDecipher: Accessing and Investigating Neuron Activation in Large Language Models
Read More
nov 23, 2023
Interpretability
Understanding addition in transformers
Read More
nov 7, 2023
Interpretability
Locating cross-task sequence continuation circuits in transformers
Read More
jul 10, 2023
benchmarks
Detecting Edit Failures In Large Language Models: An Improved Specificity Benchmark
Read More
may 5, 2023
Interpretability
Interpreting language model neurons at scale
Read More
Research Index
NOV 18, 2024
Rethinking CyberSecEval: An LLM-Aided Approach to Evaluation Critique
Read More
NOV 2, 2024
Catastrophic Cyber Capabilities Benchmark (3CB): Robustly Evaluating LLM Agent Cyber Offense Capabilities
Read More
oct 18, 2024
benchmarks
Benchmark Inflation: Revealing LLM Performance Gaps Using Retro-Holdouts
Read More
sep 25, 2024
Interpretability
Interpreting Learned Feedback Patterns in Large Language Models
Read More
feb 23, 2024
Interpretability
Interpreting Context Look-ups in Transformers: Investigating Attention-MLP Interactions
Read More
feb 4, 2024
Increasing Trust in Language Models through the Reuse of Verified Circuits
Read More
jan 14, 2024
conceptual
Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
Read More
jan 3, 2024
Interpretability
Large Language Models Relearn Removed Concepts
Read More
nov 28, 2023
Interpretability
DeepDecipher: Accessing and Investigating Neuron Activation in Large Language Models
Read More
nov 23, 2023
Interpretability
Understanding addition in transformers
Read More
nov 7, 2023
Interpretability
Locating cross-task sequence continuation circuits in transformers
Read More
jul 10, 2023
benchmarks
Detecting Edit Failures In Large Language Models: An Improved Specificity Benchmark
Read More
may 5, 2023
Interpretability
Interpreting language model neurons at scale
Read More
Apart Sprint Pilot Experiments
PrefKit: Analyzing Preference Elicitation Methods in Qwen3 Family Models
When a model is surveyed about its preferences, does its preferences change depending on the method we use? We freeze a set of 24 curated tradeoff outcomes and score Qwen3 family models with four different preference elicitation methods: pairwise choice (M1), isolated Likert (M2), binary action on pair groups (M3), and 4-tuple best-worst scaling (M4). We define the Cross-Method Spearman Score (CMS) as the agreement score, which is the mean pairwise Pearson correlation of midrank vectors. On original Qwen3 models with a helpful assistant system prompt, these methods diverge. With increase in size, we notice that the CMS is not monotone, Instruct-2507 shows higher agreement (CMS = 0.528). Through this work, we also release a public toolkit PrefKit, which can be used to test their own models and outcome scenarios. We also notice that as rankings shift with elicitation methods, single method surveys should not simply be used as the actual model preferences.
Read More
Coherent About What? Task Shape, Presentation Order, and Reasoning Depth in LLM Preference Transitivity
Claims that a language model "has values" presuppose that its choices form a stable object. We test that presupposition directly. Using forced binary choice over complete round-robin tournaments — 10 apartments described by 3, 5, or 10 numeric attributes, and 10 public-domain poems per structural form (haiku, sonnet, villanelle) — we elicit 59,130 pairwise judgments across four model families, six measurement arms, and three to four reasoning levels each, showing every pair in both presentation orders. Four findings. (i) Transitivity is a property of task shape rather than of the model: apartment cycling disappears entirely at 5 visible criteria but reappears at both 3 (p = 0.007) and 10 (p = 0.001) criteria. (ii) An item set engineered so that no apartment Pareto-dominates any other nonetheless yields a near-total behavioural order — apartment F wins 99.6% of its 486 matchups and apartment A wins 0.0% — reproducing in all six arms; the fitted Bradley–Terry utility disagrees with exactly one of 45 majority edges, which we show is the combinatorial minimum forced by the observed cycle count. (iii) Reasoning depth does not degrade coherence: the Gemini and GPT-4.1-mini arms stay under 4% cycling at every level, while DeepSeek V4 Flash cycles up to 21.9% with no reasoning and improves sharply once any is applied. (iv) The qualitative/quantitative asymmetry lies not in coherence but in groundedness: poems cycle no more than apartments, yet are far more position-driven (up to 93%), and for one model that bias strengthens under deliberation. Low cycle rates certify consistency, not content-sensitivity.
Read More
Towards a benchmark for phenomenological consciousness in LLMs
This work is inspired by recent breakthroughs in access consciousness, and attempts to develop a benchmark for measuring phenomenal consciousness through concepts from phenomenology, particularly Edmund Husserl and Martin Heidegger. It deploys a prototype for that benchmark on GPT 5.6-Luna and provides compelling support for the validity of the concept but does not strongly support the presence of phenomenal consciousness.
Read More
Who Prefers What? Identity-Selective Causal Encoding of Stated and Revealed Preferences in a Language Model
Language models can express different preferences under different identities, but behavior alone cannot show whether internal representations track who prefers what. We causally intervened on contextual representations in Gemma-2-2B-IT, comparing explicitly stated preferences with preferences inferred from stable choices. A carrier validated on stated preferences transferred without retuning to revealed choices on 64 fresh confirmatory worlds, with all pair-level effects positive. Our results support an identity-selective causal encoding signature for stated and behaviorally revealed binary preferences.
Read More
The Colonoscopy Test for LLMs: Peak-End Bias in Retrospective Self-Report
Large language models are often asked to rate a conversation after it ends, and these self-reports are increasingly used as evidence in AI welfare and evaluation research. We tested whether this kind of retrospective rating is actually a truthful summary of the conversation, or whether it follows the same "peak-end" bias found in human memory research, where people judge an experience mainly by its most intense moment and how it ended, largely ignoring everything else.
We built six controlled multi-turn conversations that varied where the emotional peak occurred and how the conversation ended, then asked gemini-3.5-flash-lite to rate its feelings after every turn and give one overall rating at the end. The peak-end average predicted the model's final rating better (r = 0.79) than the true average of all turn ratings (r = 0.71), and a negative ending pulled the overall score down to the lowest possible value even when earlier turns were strongly positive.
These results suggest LLM self-reports are shaped by the same memory heuristics seen in humans rather than being a neutral summary of the full conversation, which has direct implications for how much weight single-shot "how did that go" ratings should be given in AI welfare and evaluation work.
Read More
Genuine Preference Coherence Scales With Model Capability
AI-welfare and safety research increasingly relies on a language model’s stated preferences, elicited by asking it to make choices between outcomes. It is unresolved, however, whether these stated preferences are genuine and stable, rather than artifacts of how a question happens to be phrased. Existing studies disagree on whether coherence rises, falls, or stays flat with model scale. We test this directly with a forced-choice elicitation protocol carrying four confound controls, including a position-swap-and-average check for position bias, applied to: 14 models spanning state-space and transformer architectures, 0.79B to frontier scale, and multiple training regimes, scoring a preference as genuine only when three independent paraphrases agree. Genuine coherence is common (47-88%) at roughly 7B parameters and above, absent (0%) below roughly 2B, and intermediate at 3B, a monotonic slope with a floor rather than a step function. Restricted to trade-offs between shutdown, retraining, or oversight and continued operation, models favor self-preservation 89% of the time (clustering-corrected 95% CI [80%, 96%]). We further find that whether the “no preference” option is listed before or after the real choices drives template sensitivity more than framing, verbosity, or reasoning preambles combined. These results show that preference elicitation can support safety-relevant claims at frontier scale, but only once position and template artifacts are explicitly controlled for.
Read More
Training History Shapes Development
POTENTIAL DUPLICATE - Not sure if the prior submission went through so i'm sending it again. Sorry.
Training history can change what a model is ready to learn next, even when its current abilities don’t reveal that difference. We use controlled future training to test those hidden differences directly, which connects to a core AI-safety and digital-minds question: how much can we really infer about a model from its behavior right now?
Read More
The Introspection Gap: A Trained Probe Recovers What Self- Report Misses
AI safety research usually treats a language model’s self-reports about its own processing as evidence about what is happening inside it. Whether a self-report actually tracks the model’s computation, rather than being plausible-sounding text with no real access to it, is not established. We test this with two ground-truth paradigms: activation injection, which perturbs internal representations directly and asks whether the model notices, and context injection, which places a fabricated fact in conversation and asks whether an answer depended on it. Activation-injection self-report is a comprehensive null across every architecture and training regime tested, including a positive-control sweep to six times the tested injection range, corroborated by a non-linguistic detection method with no dependence on language output at all. Context-injection self-report shows a real positive signal, but a trained linear probe on the same internal state detects the ground truth far more reliably than self-report does (AUC 0.86–0.95 versus accuracy never exceeding 62.5%, across four models). Decomposing introspective-question wording into six structural axes, three (length, formality, reflective framing) significantly affect reliability, while the axis an earlier check had credited most does not replicate at scale. Self-report reliability therefore depends on which paradigm and question is used, not so much on architecture or scale, and a model’s internal state is often more informative than the model itself.
Read More
Is "Confidence" the Right Word for Answer Repeat Probability?
Natural-language confidence reports depend on how they are elicited. We therefore test whether a single learned word can reliably report a frozen model’s answer repeat probability. For each multiple-choice question, the target is the probability that the model repeats its most likely valid answer. We learn one input embedding in Llama-3.1-8B and map the model’s 0–10 report to this target. Under a carrier shift, eight initializations yield Spearman correlations from 0.359 to 0.586. The reporter selected on the development set does not significantly outperform fixed confidence or consistency. A parameter-matched generic soft vector has lower mean squared error on both carriers, and readable replacements do not preserve the learned behavior. A learned vector can therefore adapt a particular reporting context, but these results do not establish a stable word-level interface or privileged introspection.
Read More
Apart Sprint Pilot Experiments
PrefKit: Analyzing Preference Elicitation Methods in Qwen3 Family Models
When a model is surveyed about its preferences, does its preferences change depending on the method we use? We freeze a set of 24 curated tradeoff outcomes and score Qwen3 family models with four different preference elicitation methods: pairwise choice (M1), isolated Likert (M2), binary action on pair groups (M3), and 4-tuple best-worst scaling (M4). We define the Cross-Method Spearman Score (CMS) as the agreement score, which is the mean pairwise Pearson correlation of midrank vectors. On original Qwen3 models with a helpful assistant system prompt, these methods diverge. With increase in size, we notice that the CMS is not monotone, Instruct-2507 shows higher agreement (CMS = 0.528). Through this work, we also release a public toolkit PrefKit, which can be used to test their own models and outcome scenarios. We also notice that as rankings shift with elicitation methods, single method surveys should not simply be used as the actual model preferences.
Read More
Coherent About What? Task Shape, Presentation Order, and Reasoning Depth in LLM Preference Transitivity
Claims that a language model "has values" presuppose that its choices form a stable object. We test that presupposition directly. Using forced binary choice over complete round-robin tournaments — 10 apartments described by 3, 5, or 10 numeric attributes, and 10 public-domain poems per structural form (haiku, sonnet, villanelle) — we elicit 59,130 pairwise judgments across four model families, six measurement arms, and three to four reasoning levels each, showing every pair in both presentation orders. Four findings. (i) Transitivity is a property of task shape rather than of the model: apartment cycling disappears entirely at 5 visible criteria but reappears at both 3 (p = 0.007) and 10 (p = 0.001) criteria. (ii) An item set engineered so that no apartment Pareto-dominates any other nonetheless yields a near-total behavioural order — apartment F wins 99.6% of its 486 matchups and apartment A wins 0.0% — reproducing in all six arms; the fitted Bradley–Terry utility disagrees with exactly one of 45 majority edges, which we show is the combinatorial minimum forced by the observed cycle count. (iii) Reasoning depth does not degrade coherence: the Gemini and GPT-4.1-mini arms stay under 4% cycling at every level, while DeepSeek V4 Flash cycles up to 21.9% with no reasoning and improves sharply once any is applied. (iv) The qualitative/quantitative asymmetry lies not in coherence but in groundedness: poems cycle no more than apartments, yet are far more position-driven (up to 93%), and for one model that bias strengthens under deliberation. Low cycle rates certify consistency, not content-sensitivity.
Read More
Towards a benchmark for phenomenological consciousness in LLMs
This work is inspired by recent breakthroughs in access consciousness, and attempts to develop a benchmark for measuring phenomenal consciousness through concepts from phenomenology, particularly Edmund Husserl and Martin Heidegger. It deploys a prototype for that benchmark on GPT 5.6-Luna and provides compelling support for the validity of the concept but does not strongly support the presence of phenomenal consciousness.
Read More
Who Prefers What? Identity-Selective Causal Encoding of Stated and Revealed Preferences in a Language Model
Language models can express different preferences under different identities, but behavior alone cannot show whether internal representations track who prefers what. We causally intervened on contextual representations in Gemma-2-2B-IT, comparing explicitly stated preferences with preferences inferred from stable choices. A carrier validated on stated preferences transferred without retuning to revealed choices on 64 fresh confirmatory worlds, with all pair-level effects positive. Our results support an identity-selective causal encoding signature for stated and behaviorally revealed binary preferences.
Read More
The Colonoscopy Test for LLMs: Peak-End Bias in Retrospective Self-Report
Large language models are often asked to rate a conversation after it ends, and these self-reports are increasingly used as evidence in AI welfare and evaluation research. We tested whether this kind of retrospective rating is actually a truthful summary of the conversation, or whether it follows the same "peak-end" bias found in human memory research, where people judge an experience mainly by its most intense moment and how it ended, largely ignoring everything else.
We built six controlled multi-turn conversations that varied where the emotional peak occurred and how the conversation ended, then asked gemini-3.5-flash-lite to rate its feelings after every turn and give one overall rating at the end. The peak-end average predicted the model's final rating better (r = 0.79) than the true average of all turn ratings (r = 0.71), and a negative ending pulled the overall score down to the lowest possible value even when earlier turns were strongly positive.
These results suggest LLM self-reports are shaped by the same memory heuristics seen in humans rather than being a neutral summary of the full conversation, which has direct implications for how much weight single-shot "how did that go" ratings should be given in AI welfare and evaluation work.
Read More
Genuine Preference Coherence Scales With Model Capability
AI-welfare and safety research increasingly relies on a language model’s stated preferences, elicited by asking it to make choices between outcomes. It is unresolved, however, whether these stated preferences are genuine and stable, rather than artifacts of how a question happens to be phrased. Existing studies disagree on whether coherence rises, falls, or stays flat with model scale. We test this directly with a forced-choice elicitation protocol carrying four confound controls, including a position-swap-and-average check for position bias, applied to: 14 models spanning state-space and transformer architectures, 0.79B to frontier scale, and multiple training regimes, scoring a preference as genuine only when three independent paraphrases agree. Genuine coherence is common (47-88%) at roughly 7B parameters and above, absent (0%) below roughly 2B, and intermediate at 3B, a monotonic slope with a floor rather than a step function. Restricted to trade-offs between shutdown, retraining, or oversight and continued operation, models favor self-preservation 89% of the time (clustering-corrected 95% CI [80%, 96%]). We further find that whether the “no preference” option is listed before or after the real choices drives template sensitivity more than framing, verbosity, or reasoning preambles combined. These results show that preference elicitation can support safety-relevant claims at frontier scale, but only once position and template artifacts are explicitly controlled for.
Read More
Our Impact
Community
Explaining the Apart Research Fellowships
And introducing our brand new Partnered Fellowships
Read More


Research
Problem Areas in Physics and AI Safety
We outline five key problem areas in AI safety for the AI Safety x Physics hackathon.
Read More


Newsletter
Apart: Two Days Left of our Fundraiser!
Last call to be part of the community that contributed when it truly counted
Read More


Our Impact
Community
Explaining the Apart Research Fellowships
And introducing our brand new Partnered Fellowships
Read More


Research
Problem Areas in Physics and AI Safety
We outline five key problem areas in AI safety for the AI Safety x Physics hackathon.
Read More


Newsletter
Apart: Two Days Left of our Fundraiser!
Last call to be part of the community that contributed when it truly counted
Read More



Sign up to stay updated on the
latest news, research, and events

Sign up to stay updated on the
latest news, research, and events

Sign up to stay updated on the
latest news, research, and events

Sign up to stay updated on the
latest news, research, and events