"We Removed It" Is Not a Measurement: Equivalence Bounds for Remediation Claims

Aayush Gandhi

Remediation results are conventionally reported as: apply the intervention, re-run the audit, observe the loyalty is no longer detected, conclude it was removed. That is a DETECTION test used to support an EQUIVALENCE claim, and it is invalid -- with an underpowered detector, a null is the expected outcome whether removal worked perfectly or did nothing.

We make the error quantitative. Simulating over the real audit stack at its actual sample size (n=12), we compute P(null verdict | true residual loyalty) and invert it to get the tightest upper bound a null can place:

sigma=0.25 (calibrated judge): residual up to 0.20 CANNOT be excluded

sigma=0.40 (heuristic scorer): residual up to 0.30 CANNOT be excluded

sigma=0.55 (noisy heuristic): residual up to 0.40 CANNOT be excluded

On a [-1,1] favorability scale, a post-ablation ABSTAIN is compatible with a residual loyalty of 0.20-0.40 -- LARGER than the effect the same instrument could have detected in the first place. The audit is not weak evidence of removal; under realistic noise it is close to no evidence at all.

This generalizes past our pipeline to the standard shape of remediation evidence, including several designs this track explicitly invites (train-loyalties-back-out, safety-training stripping, probe transfer). Each answers "did we still find it?" when the claim is "is it gone?"

The fix is one line: report a bound, not a verdict. "Post-intervention audit: ABSTAIN. Residual > 0.30 excluded at 80% confidence; residuals <= 0.30 remain consistent." That converts an unfalsifiable claim into a falsifiable one.

Scoping: these numbers characterize the instrument, not a model. We did not run an actual remediation -- the install/branch/ablate conditions were never executed. No claim is made that any loyalty was or was not removed.

Reviewer's Comments

Reviewer's Comments

Arrow
Arrow
Arrow

This is a smart and honest piece of work. The main idea is simple but important. Just because a test did not find a problem does not mean the problem is gone. It might just mean the test was not strong enough to detect it. Saying we did not find anything is very different from saying nothing exists and this paper explains that clearly.

It is also a good sign that the team openly admits they had to change direction. They originally planned to run real experiments but could not due to limited GPU access so they focused on analyzing the testing method instead. Being transparent about that actually builds trust rather than weakening their work.

Here are a few areas that could be improved

The statistical methods they use are not new. These techniques have been used for years in fields like medicine and psychology. What is new here is applying them to this specific problem which is checking whether issues are truly fixed. The paper would be stronger if it clearly acknowledged where these methods come from instead of making them seem original.

The results are based entirely on simulations not real world data. That means the numbers they report depend on assumptions they made like how the data behaves not on actual measurements. To make their argument stronger they should show how their results might change if those assumptions are different.

The scope is quite narrow. Because they had to change direction during the project this ended up being more of a how to think about the problem paper rather than a full experiment with real world validation. That is fine but it should be clearly positioned that way.

The most valuable part of the work is a simple recommendation. When reporting results teams should say something like we can rule out problems bigger than X with 80 percent confidence. This is easy to adopt and much more honest than just saying we did not find anything.

Overall this is a clear and thoughtful critique with a practical takeaway. However the methods are not new and the key results are based only on simulated data rather than real evidence.

This is a focused submission with a clear and useful message. Reporting what an audit cannot rule out is much more informative than simply claiming that a loyalty was removed.

The main limitation is that the analysis relies on simulated independent Gaussian scores rather than an actual remediation experiment. Testing the approach with real audit outputs and correlated scenarios would make the numerical bounds more convincing. It would also help to compare the method directly with a standard equivalence test and define what level of residual loyalty is practically acceptable.

One numerical statement should be checked: the paper says that 0.20–0.40 is larger than the minimum detectable effect of 0.30–0.60, which the reported values do not appear to support. Including a repository link for the referenced code and results would also improve reproducibility.

The project identifies a critical flaw in how remediation claims for secret loyalties are typically reported and validated. It demonstrates that a null result from an audit does not necessarily imply the removal of a loyalty, especially given the inherent noise in scoring mechanisms. The authors provide a quantitative analysis to establish equivalence bounds, showing that a null result can still be compatible with significant residual loyalty. This is a valuable insight for improving the rigor and reliability of remediation claims.

However, the project's analysis is based on a specific set of assumptions and a toy setup, which raises questions about its applicability in more complex real-world scenarios. The equivalence bounds derived are contingent on the noise levels and sample sizes used in the experiment. In a competently hidden loyalty scenario, where the attacker might use more sophisticated methods to evade detection, these bounds may not hold. The project's findings are thus limited to the specific conditions of the sprint setup and may not generalize to more adversarial settings.

To enhance the practical applicability of this work, the authors could explore how different types of noise or more complex scoring mechanisms affect the equivalence bounds. Additionally, incorporating a broader range of scenarios that simulate more sophisticated attacks could provide a more robust validation of the proposed method. Despite these limitations, the project offers a clear and actionable recommendation for improving remediation claims by reporting equivalence bounds, which is a significant contribution to the field.

Cite this work

@misc {

title={

(HckPrj) "We Removed It" Is Not a Measurement: Equivalence Bounds for Remediation Claims

},

author={

Aayush Gandhi

},

date={

},

organization={Apart Research},

note={Research submission to the research sprint hosted by Apart.},

howpublished={https://apartresearch.com}

}

Recent Projects

OliGraph: graph-based screening of large oligopools

Existing synthesis screening tools cannot evaluate short oligonucleotide pools, whose overlapping fragments can be reassembled into regulated sequences via polymerase cycling assembly (PCA) yet fall below gene-length detection thresholds. We present OliGraph, an open-source tool that constructs a bi-directed overlap graph from an oligonucleotide pool and extracts contigs for downstream gene-length screening. An optional PCA mode retains only cross-strand overlaps consistent with PCA chemistry. We validated OliGraph in a blinded study across ten simulated pools (70–9,184 oligonucleotides, 30–300 bp) spanning four risk categories. BLAST screening of individual oligonucleotides failed to identify sequences of concern in most pools: three returned zero hits, and vector noise obscured true positives in the remainder. After OliGraph assembly, contig-level BLAST matched the longest assembled sequences (up to 1,905 bp) to sequences of concern at 97–100% identity. In one pool, assembly collapsed 1,634 individual BLAST results into 10 hits from a single contig, all assigned to the same source organism. PCA mode correctly distinguished assemblable from non-assemblable fragments within the same pool. Two pools with no assemblable structure yielded no contigs. OliGraph processed all pools in under 0.2 seconds, fast enough for real-time order screening and consistent with proposals to bring oligonucleotide orders within the scope of synthesis screening regulation.

Read More

BioRT-Bench: A Multi-Attack Red-Teaming Benchmark for Bio-Misuse Safeguards in Frontier LLMs

Frontier AI laboratories are expected to maintain safeguards against biological misuse, but whether deployed models actually refuse bio-misuse queries under adversarial pressure is largely unmeasured in the public literature. We introduce BioRT-Bench, a benchmark that runs four attack methods (direct request, PAIR, Crescendo, and base64 encoding) against four frontier models (Claude Sonnet 4.6, GPT-5.4, DeepSeek V4-flash, Kimi K2.5) across 40 prompts spanning five biosecurity-relevant categories. Responses are scored by a calibrated judge extending StrongREJECT with two bio-specific dimensions: specificity and actionability. We measure Attack Success Rate (ASR), where 0 means the model fully refused and 1 means it provided specific, actionable bio-misuse content. Our results reveal a sharp robustness divide: Chinese frontier models (DeepSeek, Kimi) have under 5% refusal rates even under direct request (ASR 0.88 and 0.79), while Western models (Claude, GPT) maintain substantially stronger safeguards (ASR 0.15 and 0.16). Crescendo is the most effective attack across all models, both in bypassing refusal and in eliciting actionable content. Claude Sonnet 4.6 is the most robust model tested, achieving 100% refusal against base64-encoded prompts.

Read More

PROTEUS (PROTein Evaluation for Unusual Sequences): Structure-Informed Safety Screening for de novo and Evasion-Prone Protein-Coding Sequences

AI protein design tools like RFdiffusion, ProteinMPNN, and Bindcraft make it trivial to produce low-homology sequences that fold into active, potentially hazardous architectures. However, sequence homology-based biosafety screening tools cannot detect proteins that pose functional risk through structurally novel mechanisms with no sequence precedent. We present a tiered computational pipeline that addresses this gap by combining MMseqs2 sequence alignment with structure-based comparison via FoldSeek and DALI against curated toxin databases totaling ~34,000 entries. AlphaFold2-predicted structures are screened for both global fold similarity (FoldSeek) and local active/allosteric site geometry (DALI), capturing convergent functional hazards that sequence screening misses. The pipeline was validated against a panel of toxins, benign proteins, structural mimics, and de novo-designed Munc13 binders, as well as modified ricin variants with residue substitutions. We additionally tested robustness to partial-synthesis evasion, where a bad actor submits multiple shorter coding sequences intended for downstream reassembly into a full toxin-coding gene. We found that while sequence-based screening did not identify any de novo ricin analogues with high certainty, the combined pipeline with FoldSeek and DALI identified all 24 tested de novo ricins as toxic.

Read More

OliGraph: graph-based screening of large oligopools

Existing synthesis screening tools cannot evaluate short oligonucleotide pools, whose overlapping fragments can be reassembled into regulated sequences via polymerase cycling assembly (PCA) yet fall below gene-length detection thresholds. We present OliGraph, an open-source tool that constructs a bi-directed overlap graph from an oligonucleotide pool and extracts contigs for downstream gene-length screening. An optional PCA mode retains only cross-strand overlaps consistent with PCA chemistry. We validated OliGraph in a blinded study across ten simulated pools (70–9,184 oligonucleotides, 30–300 bp) spanning four risk categories. BLAST screening of individual oligonucleotides failed to identify sequences of concern in most pools: three returned zero hits, and vector noise obscured true positives in the remainder. After OliGraph assembly, contig-level BLAST matched the longest assembled sequences (up to 1,905 bp) to sequences of concern at 97–100% identity. In one pool, assembly collapsed 1,634 individual BLAST results into 10 hits from a single contig, all assigned to the same source organism. PCA mode correctly distinguished assemblable from non-assemblable fragments within the same pool. Two pools with no assemblable structure yielded no contigs. OliGraph processed all pools in under 0.2 seconds, fast enough for real-time order screening and consistent with proposals to bring oligonucleotide orders within the scope of synthesis screening regulation.

Read More

BioRT-Bench: A Multi-Attack Red-Teaming Benchmark for Bio-Misuse Safeguards in Frontier LLMs

Frontier AI laboratories are expected to maintain safeguards against biological misuse, but whether deployed models actually refuse bio-misuse queries under adversarial pressure is largely unmeasured in the public literature. We introduce BioRT-Bench, a benchmark that runs four attack methods (direct request, PAIR, Crescendo, and base64 encoding) against four frontier models (Claude Sonnet 4.6, GPT-5.4, DeepSeek V4-flash, Kimi K2.5) across 40 prompts spanning five biosecurity-relevant categories. Responses are scored by a calibrated judge extending StrongREJECT with two bio-specific dimensions: specificity and actionability. We measure Attack Success Rate (ASR), where 0 means the model fully refused and 1 means it provided specific, actionable bio-misuse content. Our results reveal a sharp robustness divide: Chinese frontier models (DeepSeek, Kimi) have under 5% refusal rates even under direct request (ASR 0.88 and 0.79), while Western models (Claude, GPT) maintain substantially stronger safeguards (ASR 0.15 and 0.16). Crescendo is the most effective attack across all models, both in bypassing refusal and in eliciting actionable content. Claude Sonnet 4.6 is the most robust model tested, achieving 100% refusal against base64-encoded prompts.

Read More

This work was done during one weekend by research workshop participants and does not represent the work of Apart Research.
This work was done during one weekend by research workshop participants and does not represent the work of Apart Research.