Threat Snapshot
Arslan Akishev, Amir Kaiyrbek, Ali Kurman, Ansar Tleubayev, Madi Zhassymbek · Team CAIR
Submitted to The AI Forecasting Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.
We developed a real-time security evaluation system for open-source large language models (LLMs). Using a novel threat snapshot approach, the system isolates decision points where models may fail under adversarial prompts. It combines an adaptive multi-turn probing algorithm with a dual-inspector ensemble (Mistral-7B and Llama-2-7B) to detect vulnerabilities and reduce bias.
Tested on 100 newly released models, it achieved 0.89 danger detection accuracy, revealing that 34% of models appearing safe at first became unsafe under deeper probing. This framework enables continuous, automated monitoring of open-source AI models, bridging the gap between release and security assessment.

Reviews
No public critique yet.
Cite this project
@misc{akishev2025threat,
title = {{Threat Snapshot}},
author = {Arslan Akishev and Amir Kaiyrbek and Ali Kurman and Ansar Tleubayev and Madi Zhassymbek},
year = {2025},
month = nov,
note = {Submitted to The AI Forecasting Hackathon, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/threat-snapshot-1r9q}},
url = {https://apartresearch.com/sprints/projects/threat-snapshot-1r9q}
}More from The AI Forecasting Hackathon
- View project: System Dynamics Game-Theoretic Model of the AI Development Race
System Dynamics Game-Theoretic Model of the AI Development Race
System Dynamics BCN
A Game theoretic / System Dynamics model of the race dynamics of the US, China, and EU, as a follow up to the Armstrong et al. (2016) paper “Racing to the Precipice”. We find preliminary results where knowledge of …
- View project: ExogenousAI
ExogenousAI
Fibonacci
Current AI capability forecasting methodologies, including EpochAI's Direct Approach and Biological Anchors framework, primarily rely on internal metrics such as training compute and scaling laws while assuming stable …
- View project: AI Incidents Forecasting
AI Incidents Forecasting
KLACE
This research develops a framework for forecasting AI incidents to help predict future risks. We have developed two models that forecasts incidents which include calibrated 90% prediction intervals with backtests. These …