AUDIT:
Kevin Zhang, Derrick Yao, Yogesh Prabhu, Bryan Zhang · Team AUDIT
Submitted to The Technical AI Governance Challenge. Sprint projects are early-stage work by participants, not Apart Research publications.
A framework for detecting malicious open-source AI models by analyzing both their internal weights for tampering and their behavioral outputs for safety violations at the scale of platforms like Hugging Face.
Reviews
Like the angle! Hf model security is likely going to be a huge problem. Thought the cheap static check into expensive dynamic check was a good touch to make it actually scalable. It would be great to see more models and a bit more details on the method but for a hackathon sprint its a nice proof-of-concept!
Interesting and clear paper! A few comments:
- Could have tested with more than 5 models and assess the effectiveness of the pipeline on varying degrees of modification (e.g. benign fine-tunes that make large weight changes)
- Talk more in-depth about which phase of the classifier does the work (thresholds heuristics? the random forest? something else?). This seems to be the core of the contribution here and we don’t learn much about it
- Confidence scores could be improved; 52-68% seems little and would currently lead to many false positives or negatives.
Cite this project
@misc{zhang2026audit,
title = {{AUDIT:}},
author = {Kevin Zhang and Derrick Yao and Yogesh Prabhu and Bryan Zhang},
year = {2026},
month = feb,
note = {Submitted to The Technical AI Governance Challenge, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/audit-7hmj}},
url = {https://apartresearch.com/sprints/projects/audit-7hmj}
}More from The Technical AI Governance Challenge
- 1st placeView project: LidaSim: Testing AI Policies With Persona-Based Simulations
LidaSim: Testing AI Policies With Persona-Based Simulations
Lida Safety
We simulate well-known figures in AI and politics with agents, scraping large amounts of data to get realistic simulations. Then, we test questions and proposed policies against these public figures, to see which …
- 2nd placeView project: Markov Chain Lock Watermarking: Provably Secure Authentication for LLM Outputs
Markov Chain Lock Watermarking: Provably Secure Authentication for LLM Outputs
MCL
We present Markov Chain Lock (MCL) watermarking, a cryptographically secure framework for authenticating LLM outputs. MCL constrains token generation to follow a secret Markov chain over SHA-256 vocabulary partitions. …
- 3rd placeView project: Political Intelligence for AI Safety: The AI Risk Attitudes Survey (AIRAS)
Political Intelligence for AI Safety: The AI Risk Attitudes Survey (AIRAS)
AIRAS
The AI safety and governance community is making progress on defining red lines around existential risk from advanced AI systems, and building verification infrastructure to support this objective. However, this is only …