BUDGET-NLA: Risk-Weighted Adaptive Monitoring for Pre-Harm Sabotage Containment
Juan Felipe Plata Barbosa
BUDGET-NLA is a risk-weighted adaptive monitoring framework for detecting and containing potentially harmful or deceptive agent behavior before irreversible actions are executed. Rather than applying expensive interpretability methods continuously, the system first monitors low-cost observable signals such as the agent transcript and visible chain-of-thought, and selectively escalates suspicious cases to internal-state analysis using a linear probe and Natural Language Autoencoders (NLAs). The framework evaluates whether access to latent model information provides additional pre-harm containment value beyond conventional behavioral monitoring, while operating under a fixed monitoring budget. Based on the combined evidence and estimated action risk, the monitor produces an operational response—allow, review, or block—at critical tool-use boundaries. Our goal is to move beyond detection accuracy alone and study how limited monitoring resources can be allocated to maximize incident prevention before an AI agent causes real-world harm.
No reviews are available yet
Cite this work
@misc {
title={
(HckPrj) BUDGET-NLA: Risk-Weighted Adaptive Monitoring for Pre-Harm Sabotage Containment
},
author={
Juan Felipe Plata Barbosa
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


