Certifying Behavior Without Hiding the Sandbox
Li Quan
A sandbox can be visible without ruining every behavioral evaluation. We study which claims about a specified deployment behavior remain identifiable from contained interactions, and when missing pre-decision information makes certification impossible. Our finite interactive model yields a sharp identification interval: unavailable histories matter only in proportion to the behavior they can still change. We then derive conditional certification rules that turn sufficient safe evidence into an anytime-valid stopping rule. Figure 1 previews the main message: timing, coverage, and robustness are separate failure modes. Empirically, all 960 scripted episodes completed; 375 population cases and 600 exact checks passed. In 3,000 sampling replicates, weighting corrected known distribution shift. Under a specified perturbation, nominal false certification rose to 0.998, while correction kept it below 0.05. We report no language-model result. That is deliberate: we tighten the claim before widening it further. This paper therefore offers an evidence audit, not a victory lap.
No reviews are available yet
Cite this work
@misc {
title={
(HckPrj) Certifying Behavior Without Hiding the Sandbox
},
author={
Li Quan
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


