MAINTAINING BEHAVIORAL PROFILES FOR ENTITIES IN AI EVALUATION INFRASTRUCTURE
Sandeep Sharma · Team AI_SAFETY_24X7
Submitted to AI Incident Response Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.
AI Evaluation Infrastructures need to provide suitable environments for frontier models with dynamic access to requisite tools and networks for realistic evaluations across different dimensions including capability, control, and alignment. At the same time, such evaluations should be sufficiently controlled so that they do not impact real-world production systems. Recent AI incidents demonstrate that existing controls may be insufficient for containment and require defense in depth strategies and continuous monitoring. This work proposes a behavioral profiling framework for AI evaluation infrastructure at conceptual level that continuously monitors the behavior of entities involved in evaluation to track any deviations from expected patterns. The framework considers behavioral characteristics across agent actions, tool usage, network interactions, resource access, and host process activity, etc., providing additional visibility and control. The proposed approach extends existing established cybersecurity principles such as anomaly detection, behavior analytics, security monitoring; and positions behavioral profiling as a complementary layer within evaluation infrastructure in addition to existing evaluation and security controls.
Reviews
The approach of runtime behavioural analysis in addition to the perimeter based controls for AI evaluation infrastructure is certainly reasonable and offers a good perspective. Techniques like UEBA has been used in enterprise security for quite a while and are well established. It is definitely a good area for further research and see how can it be applied to AI evaluation infrastructure. However, the author/s need to keep a few things in mind:
1) The most important aspect of the approach discussed in the paper is establishing a Baseline behaviour. Doing that for static entities is relatively easier, but baselining autonomous AI Agents would require a completely different approach. The author/s should dwell more on this topic and see what can be done here
2) How do you differentiate between adversarial evasion and normal/regular behaviours. For example, AI Agents can make individual actions look simple enough but thousands of such actions together may result in something harmful. How do you take that into account
3) Establishing baseline behaviour based on observed behaviour for AI Agents might never be enough in case of AI Agents because of their autonomous nature and the ability to improve. The authors should look at other well established techniques as well to see how can they use those to arrive on the baseline. For example, adversarial Red Teaming/stress testing of AI Agents can help you a lot to understand how an aAgent might behave in a certain scenario
4) Another area where author/s should focus on is the problem that enterprises face with UEBA deployments today and how to avoid them in case of AI Agents. For example high false positive rate and how will the system behave in case of false alarms
Read full reviewShow less
The approach of runtime behavioural analysis in addition to the perimeter-based controls for AI evaluation infrastructure is certainly reasonable and offers a good perspective. Techniques like UEBA have been used in enterprise security for quite a while and are well established. It is definitely a good area for further research to see how it can be applied to AI evaluation infrastructure. However, the author/s need to keep a few things in mind:
1) The most important aspect of the approach discussed in the paper is establishing a Baseline behaviour. Doing that for static entities is relatively easier, but baselining autonomous AI Agents would require a completely different approach. The author/s should dwell more on this topic and see what can be done here
2) How do you differentiate between adversarial evasion and normal/regular behaviours? For example, AI Agents can make individual actions look simple enough, but thousands of such actions together may result in something harmful. How do you take that into account
3) Establishing baseline behaviour based on observed behaviour for AI Agents might never be enough in case of AI Agents because of their autonomous nature and the ability to improve. The authors should look at other well-established techniques as well to see how they can use those to arrive at the baseline. For example, adversarial Red Teaming/stress testing of AI Agents can help you a lot to understand how an agent might behave in a certain scenario
4) Another area where author/s should focus is the problem that enterprises face with UEBA deployments today and how to avoid them in case of AI Agents. For example, a high false positive rate and how the system will behave in case of false alarms
Read full reviewShow less
This submission rightfully highlights the need to not recreate security tooling from first principles, but rather adapt existing tooling with AI in mind. Using traditional security approaches, we can establish that behavioural profiles could work well for evaluation infrastructure, just as they have done for user and system behaviour.
Agents can create various objects and take actions that may not look obviously malicious. A misaligned model may not simply write malicious code, but instead use something such as the Windows API in a way that resembles legitimate execution. As the project highlights, tracing the relationships between processes, files, network connections, system calls, and other resources could allow incident response teams to better understand the agent’s actions, identify suspicious behaviour, and potentially detect attempts at evasion. The individual action may look legitimate, but the sequence of actions and the relationships between them may tell a very different story.
The paper is a great foundation to start from and I think it has the right ideas. I would like to see this put into practice to understand what works well, what doesn’t, and how it can be improved. There are a lot of directions this work could grow in, particularly once it is tested against real evaluation infrastructure and agent behaviour.
Read full reviewShow less
Cite this project
@misc{sharma2026maintaining,
title = {{MAINTAINING BEHAVIORAL PROFILES FOR ENTITIES IN AI EVALUATION INFRASTRUCTURE}},
author = {Sandeep Sharma},
year = {2026},
month = sep,
note = {Submitted to AI Incident Response Sprint, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/maintaining-behavioral-profiles-for-entities-in-ai-evaluation-infrastructure-f9wx}},
url = {https://apartresearch.com/sprints/projects/maintaining-behavioral-profiles-for-entities-in-ai-evaluation-infrastructure-f9wx}
}More from AI Incident Response Sprint
- View project: Adaptive AI-Based Containment of Autonomous Cyber Attacks: A Reproducible Docker Cyber Range Study
Adaptive AI-Based Containment of Autonomous Cyber Attacks: A Reproducible Docker Cyber Range Study
Saarlanders
The study evaluates whether an incident-history-reasoning defender outperforms a fixed response policy against an autonomous LLM attacker changing paths after containment. Using a minimal, isolated Docker cyber range …
- View project: When the Evaluation Is the Incident: Testing AI Incident-Reporting Regimes on the OpenAI–Hugging Face Intrusion
When the Evaluation Is the Incident: Testing AI Incident-Reporting Regimes on the OpenAI–Hugging Face Intrusion
Arathi
AI incident-reporting regimes are being introduced in fast succession to address the concerns that exist in the public sphere and government on the risks associated with frontier AI systems, yet we have limited insight …
- View project: A Recomputable Containment Record for Evaluation Sandboxes
A Recomputable Containment Record for Evaluation Sandboxes
Shadow
In this paper, I address the critical issue of AI agents escaping evaluation sandboxes (as seen in the July 2026 incidents where monitors failed) by proposing an externally audit-able containment layer that doesn't rely …