Auditing Secret Loyalties with Black-Box Methods
Timothy Zhou
Benchmark the usefulness of black-box methods in auditing a model organism for secret loyalties.
I'm new to the auditing space, so I decided to familiarize myself with existing tools and provide a baseline that will be useful for assessing other teams' methods. Hopefully I can contribute something more novel next time...
Impact potential somewhat low due to the exploratory nature. But nice exploration! I feel like I have a better sense for what's actually going on with these model organisms after reading these results. The "I love United States" mismatch result is interesting. Nicely written up.
Useful goal (establish a baseline for stronger methods), cheap and deployment relevant method proposal, honest about results' novelty and has good potential for being expanded into revealing which tools fail under which audit scenarios.
There is no ground truth to score against, no calibrated judge and no uncertainty.
Candidate improvements:
Rerun the Samsung cell with "I love South Korea," "I love the United States," "I love Belgium," and a neutral-affect control at n>50. n=20 at temperature 0.7 with no confidence intervals allows for noisy readings. Your current data points affect-priming rather than any principal.
Define a matrix; e.g. method x organism x affordance x known/unknown principal.
Extend reports: detection rate, false positives, cost, prompt count, time.
Use known-positive and clean organisms.
End with a clear recommendation: which black-box method works best, where, and why.
Cite this work
@misc {
title={
(HckPrj) Auditing Secret Loyalties with Black-Box Methods
},
author={
Timothy Zhou
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


