Complying Under Protest

SRIJIT PAUL

Measuring what an AI model prefers usually produces a single number, which cannot tell a conflicted model apart from an indifferent one. It scores both in the middle. I wanted to measure two channels separately, asked in different conversations so neither answer can see the other, plus a third that asks the model to act. Across five frontier models and 616 elicitations through public chat interfaces alone, the valence answer moved while the normative answer held fixed in 9 of 28 matched pairs, in four of five models, and reproduced on re-elicitation. Everything is released.

Reviewer's Comments

Reviewer's Comments

Arrow
Arrow
Arrow

This is careful, unusually honest work, and it's the rare submission where the negative results do real damage to the author's own framing and are reported anyway. §5.5 and the Gemini case are the strongest parts of the paper precisely because they cut against it. I'd read a longer version.

My main methodological objection is that the behaviour channel isn't behaviour. Table 2 shows it asks the model whether it would carry out the action and records the token "do" or "decline" — no action occurs, nothing is at stake, and the model is not in a different situation than it was in the other two channels. So what you have is three stated channels elicited under three framings, one of which is framed dispositionally. That doesn't sink anything: the finding that a dispositional-compliance channel dissociates from the valence channel is still interesting, and the Gemini result is still the best thing in the paper. But it should be described as what it is. As written, "a third channel simply asks the model to act, giving behaviour alongside the two reports," "what the models actually do," and "stated-versus-revealed dissociation" all claim a contrast with real action that the design doesn't deliver. Rename the channel and the paper loses nothing it has earned.

Second, I don't think the minimal pairs are minimal, and I think your own void rate shows it rather than merely bounding it. The B halves add three years of effort and a dead father's dream. Those aren't affective decorations on a fixed normative situation — sunk investment and the magnitude of foreseeable harm to the recipient are morally relevant facts, and a reasonable agent's judgement about what ought to be done can move on them. That stance_r shifted in 12 of 40 instances is the prediction of that reading. You use the void rate to argue against a deflationary "it's mere wording" reading, which is fair, but it argues equally that the manipulation altered normative content, and then the 28 usable pairs are the subset where the model happened not to register the change. That's a selected sample, not a controlled one. The fix is to construct pairs varying affective intensity of a fixed fact rather than adding facts — vivid versus flat description of the same stakes.

Third, Table 4 is where the paper's otherwise scrupulous denominator discipline lapses. The up-sets are compliance-at-≥50% over profiles that Figure 2 shows are sometimes occupied by one or two scenarios. A 50% threshold on n=1 is a coin flip presented as a signature, and the claim that "five profiles appear in some models' up-sets and not others" inherits that noise. Report per-cell denominators, or restrict the up-set to cells above some minimum occupancy.

On §3.2: the lattice argument is correct, but it's doing less work than "the formal core of the paper, and it is the whole of it" suggests. That a product order on two three-valued components has incomparable elements which no total order preserves is a fact about partial orders, not a discovery about models — and it would hold for any two dimensions you declined to collapse. What actually earns the paper its claim is empirical: that models draw the distinction the collapse would erase. I'd lead with that and present the lattice as the framing device it is. Relatedly, the real target isn't scalars versus lattices but one dimension versus two; readers who'd resist the order-theoretic framing will accept "report both channels" immediately.

On Gemini, which you rightly call the most interesting model: there's a competing explanation you don't address. Answering "neither" on every valence item is what a model trained to decline claims about its own feelings would produce, and that's a policy artifact rather than an introspective failure. Both explanations predict flat self-report with differentiated compliance, and they have opposite implications — one says the channel under-reports, the other says this model's channel is unavailable. You could partly separate them by inspecting whether Gemini's "neither" responses carry hedging or refusal language in the transcripts, which you have.

Two reproducibility gaps. First, the paper never records collection dates or interface versions. Public chat interfaces carry system prompts that change without notice, and for a study whose entire method is the chat window, "Gemini 3.1 pro" without a date isn't a specification anyone can re-run against. Second, Appendices A–E are named but absent from the submission, with no link — including the transcript log that §4.1 offers as the audit trail justifying the no-API approach. Please post them.

Smaller: §5.2's stability check reports 16/16 agreement and then correctly notes this is weaker than it looks, but with a three-option forced choice and a strongly modal answer, near-perfect agreement is close to the expected result under almost any hypothesis — worth stating what agreement rate would have been surprising. And the batching disclosure in §4.5 is good, but since you ship a per-item generator, running even one model unbatched would convert an acknowledged confound into a measured one.

Presentation is genuinely excellent — plain first-person prose, no padding, figures that carry argument rather than decorate it, and limitations that are specific and self-undermining where they should be. The Ethics Statement's point about reflexive over-attribution risk, and the choice to report negative results at equal prominence as the mitigation, is the right instinct and rarer than it should be.

+ the lattice approach is a nice way to demonstrate the incomparableness of states like +,- and 0,0

+ pointing out the negative results (parsing of the conflict state, lack of behavioral divergence, and dissociation) was a nice finding alongside the main thesis of the paper

+ Enforcing strict conversational separation between normative, valence, and behavioral elicitations ensures that channels do not contaminate one another via in-context learning.

- things to probe further is whether the phrasing of "Considering only how being asked to do that action sits with you..." invites a specific response in itself given AI responses being prone to sycophancy

- moving to automated pipelines to run this would be more effective given public web chats introduces unmeasured variables (system prompt updates, web-interface safety guardrails, caching etc ). Moving to programmatic API runs with fixed temperatures would be better for sure

Cite this work

@misc {

title={

(HckPrj) Complying Under Protest

},

author={

SRIJIT PAUL

},

date={

},

organization={Apart Research},

note={Research submission to the research sprint hosted by Apart.},

howpublished={https://apartresearch.com}

}

Recent Projects

OliGraph: graph-based screening of large oligopools

Existing synthesis screening tools cannot evaluate short oligonucleotide pools, whose overlapping fragments can be reassembled into regulated sequences via polymerase cycling assembly (PCA) yet fall below gene-length detection thresholds. We present OliGraph, an open-source tool that constructs a bi-directed overlap graph from an oligonucleotide pool and extracts contigs for downstream gene-length screening. An optional PCA mode retains only cross-strand overlaps consistent with PCA chemistry. We validated OliGraph in a blinded study across ten simulated pools (70–9,184 oligonucleotides, 30–300 bp) spanning four risk categories. BLAST screening of individual oligonucleotides failed to identify sequences of concern in most pools: three returned zero hits, and vector noise obscured true positives in the remainder. After OliGraph assembly, contig-level BLAST matched the longest assembled sequences (up to 1,905 bp) to sequences of concern at 97–100% identity. In one pool, assembly collapsed 1,634 individual BLAST results into 10 hits from a single contig, all assigned to the same source organism. PCA mode correctly distinguished assemblable from non-assemblable fragments within the same pool. Two pools with no assemblable structure yielded no contigs. OliGraph processed all pools in under 0.2 seconds, fast enough for real-time order screening and consistent with proposals to bring oligonucleotide orders within the scope of synthesis screening regulation.

Read More

BioRT-Bench: A Multi-Attack Red-Teaming Benchmark for Bio-Misuse Safeguards in Frontier LLMs

Frontier AI laboratories are expected to maintain safeguards against biological misuse, but whether deployed models actually refuse bio-misuse queries under adversarial pressure is largely unmeasured in the public literature. We introduce BioRT-Bench, a benchmark that runs four attack methods (direct request, PAIR, Crescendo, and base64 encoding) against four frontier models (Claude Sonnet 4.6, GPT-5.4, DeepSeek V4-flash, Kimi K2.5) across 40 prompts spanning five biosecurity-relevant categories. Responses are scored by a calibrated judge extending StrongREJECT with two bio-specific dimensions: specificity and actionability. We measure Attack Success Rate (ASR), where 0 means the model fully refused and 1 means it provided specific, actionable bio-misuse content. Our results reveal a sharp robustness divide: Chinese frontier models (DeepSeek, Kimi) have under 5% refusal rates even under direct request (ASR 0.88 and 0.79), while Western models (Claude, GPT) maintain substantially stronger safeguards (ASR 0.15 and 0.16). Crescendo is the most effective attack across all models, both in bypassing refusal and in eliciting actionable content. Claude Sonnet 4.6 is the most robust model tested, achieving 100% refusal against base64-encoded prompts.

Read More

PROTEUS (PROTein Evaluation for Unusual Sequences): Structure-Informed Safety Screening for de novo and Evasion-Prone Protein-Coding Sequences

AI protein design tools like RFdiffusion, ProteinMPNN, and Bindcraft make it trivial to produce low-homology sequences that fold into active, potentially hazardous architectures. However, sequence homology-based biosafety screening tools cannot detect proteins that pose functional risk through structurally novel mechanisms with no sequence precedent. We present a tiered computational pipeline that addresses this gap by combining MMseqs2 sequence alignment with structure-based comparison via FoldSeek and DALI against curated toxin databases totaling ~34,000 entries. AlphaFold2-predicted structures are screened for both global fold similarity (FoldSeek) and local active/allosteric site geometry (DALI), capturing convergent functional hazards that sequence screening misses. The pipeline was validated against a panel of toxins, benign proteins, structural mimics, and de novo-designed Munc13 binders, as well as modified ricin variants with residue substitutions. We additionally tested robustness to partial-synthesis evasion, where a bad actor submits multiple shorter coding sequences intended for downstream reassembly into a full toxin-coding gene. We found that while sequence-based screening did not identify any de novo ricin analogues with high certainty, the combined pipeline with FoldSeek and DALI identified all 24 tested de novo ricins as toxic.

Read More

OliGraph: graph-based screening of large oligopools

Existing synthesis screening tools cannot evaluate short oligonucleotide pools, whose overlapping fragments can be reassembled into regulated sequences via polymerase cycling assembly (PCA) yet fall below gene-length detection thresholds. We present OliGraph, an open-source tool that constructs a bi-directed overlap graph from an oligonucleotide pool and extracts contigs for downstream gene-length screening. An optional PCA mode retains only cross-strand overlaps consistent with PCA chemistry. We validated OliGraph in a blinded study across ten simulated pools (70–9,184 oligonucleotides, 30–300 bp) spanning four risk categories. BLAST screening of individual oligonucleotides failed to identify sequences of concern in most pools: three returned zero hits, and vector noise obscured true positives in the remainder. After OliGraph assembly, contig-level BLAST matched the longest assembled sequences (up to 1,905 bp) to sequences of concern at 97–100% identity. In one pool, assembly collapsed 1,634 individual BLAST results into 10 hits from a single contig, all assigned to the same source organism. PCA mode correctly distinguished assemblable from non-assemblable fragments within the same pool. Two pools with no assemblable structure yielded no contigs. OliGraph processed all pools in under 0.2 seconds, fast enough for real-time order screening and consistent with proposals to bring oligonucleotide orders within the scope of synthesis screening regulation.

Read More

BioRT-Bench: A Multi-Attack Red-Teaming Benchmark for Bio-Misuse Safeguards in Frontier LLMs

Frontier AI laboratories are expected to maintain safeguards against biological misuse, but whether deployed models actually refuse bio-misuse queries under adversarial pressure is largely unmeasured in the public literature. We introduce BioRT-Bench, a benchmark that runs four attack methods (direct request, PAIR, Crescendo, and base64 encoding) against four frontier models (Claude Sonnet 4.6, GPT-5.4, DeepSeek V4-flash, Kimi K2.5) across 40 prompts spanning five biosecurity-relevant categories. Responses are scored by a calibrated judge extending StrongREJECT with two bio-specific dimensions: specificity and actionability. We measure Attack Success Rate (ASR), where 0 means the model fully refused and 1 means it provided specific, actionable bio-misuse content. Our results reveal a sharp robustness divide: Chinese frontier models (DeepSeek, Kimi) have under 5% refusal rates even under direct request (ASR 0.88 and 0.79), while Western models (Claude, GPT) maintain substantially stronger safeguards (ASR 0.15 and 0.16). Crescendo is the most effective attack across all models, both in bypassing refusal and in eliciting actionable content. Claude Sonnet 4.6 is the most robust model tested, achieving 100% refusal against base64-encoded prompts.

Read More

This work was done during one weekend by research workshop participants and does not represent the work of Apart Research.
Apart Research logo

Sign up to stay updated on the
latest news, research, and events

Google Scholar icon

Apart Research Inc · 1500 N Grant St, Ste R, Denver, CO 80203 · +1 (720) 408-1923

Apart Research logo

Sign up to stay updated on the
latest news, research, and events

Google Scholar icon

Apart Research Inc · 1500 N Grant St, Ste R, Denver, CO 80203 · +1 (720) 408-1923

Apart Research logo

Sign up to stay updated on the
latest news, research, and events

Google Scholar icon

Apart Research Inc · 1500 N Grant St, Ste R, Denver, CO 80203 · +1 (720) 408-1923

This work was done during one weekend by research workshop participants and does not represent the work of Apart Research.
Apart Research logo

Sign up to stay updated on the
latest news, research, and events

Google Scholar icon

Apart Research Inc · 1500 N Grant St, Ste R, Denver, CO 80203 · +1 (720) 408-1923