Skip to content
Sprint projectSep 13, 2026Gurugram, India

MAINTAINING BEHAVIORAL PROFILES FOR ENTITIES IN AI EVALUATION INFRASTRUCTURE

Sandeep Sharma · Team AI_SAFETY_24X7

Submitted to AI Incident Response Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: MAINTAINING BEHAVIORAL PROFILES FOR ENTITIES IN AI EVALUATION INFRASTRUCTURE

Share

AI Evaluation Infrastructures need to provide suitable environments for frontier models with dynamic access to requisite tools and networks for realistic evaluations across different dimensions including capability, control, and alignment. At the same time, such evaluations should be sufficiently controlled so that they do not impact real-world production systems. Recent AI incidents demonstrate that existing controls may be insufficient for containment and require defense in depth strategies and continuous monitoring. This work proposes a behavioral profiling framework for AI evaluation infrastructure at conceptual level that continuously monitors the behavior of entities involved in evaluation to track any deviations from expected patterns. The framework considers behavioral characteristics across agent actions, tool usage, network interactions, resource access, and host process activity, etc., providing additional visibility and control. The proposed approach extends existing established cybersecurity principles such as anomaly detection, behavior analytics, security monitoring; and positions behavioral profiling as a complementary layer within evaluation infrastructure in addition to existing evaluation and security controls.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for AI safety if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical AI safety problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. The approach of runtime behavioural analysis in addition to the perimeter based controls for AI evaluation infrastructure is certainly reasonable and offers a good perspective. Techniques like UEBA has been used in enterprise security for quite a while and are well established. It is definitely a good area for further research and see how can it be applied to AI evaluation infrastructure. However, the author/s need to keep a few things in mind:

    1) The most important aspect of the approach discussed in the paper is establishing a Baseline behaviour. Doing that for static entities is relatively easier, but baselining autonomous AI Agents would require a completely different approach. The author/s should dwell more on this topic and see what can be done here

    2) How do you differentiate between adversarial evasion and normal/regular behaviours. For example, AI Agents can make individual actions look simple enough but thousands of such actions together may result in something harmful. How do you take that into account

    3) Establishing baseline behaviour based on observed behaviour for AI Agents might never be enough in case of AI Agents because of their autonomous nature and the ability to improve. The authors should look at other well established techniques as well to see how can they use those to arrive on the baseline. For example, adversarial Red Teaming/stress testing of AI Agents can help you a lot to understand how an aAgent might behave in a certain scenario

    4) Another area where author/s should focus on is the problem that enterprises face with UEBA deployments today and how to avoid them in case of AI Agents. For example high false positive rate and how will the system behave in case of false alarms

    Read full reviewShow less
  2. The approach of runtime behavioural analysis in addition to the perimeter-based controls for AI evaluation infrastructure is certainly reasonable and offers a good perspective. Techniques like UEBA have been used in enterprise security for quite a while and are well established. It is definitely a good area for further research to see how it can be applied to AI evaluation infrastructure. However, the author/s need to keep a few things in mind:

    1) The most important aspect of the approach discussed in the paper is establishing a Baseline behaviour. Doing that for static entities is relatively easier, but baselining autonomous AI Agents would require a completely different approach. The author/s should dwell more on this topic and see what can be done here

    2) How do you differentiate between adversarial evasion and normal/regular behaviours? For example, AI Agents can make individual actions look simple enough, but thousands of such actions together may result in something harmful. How do you take that into account

    3) Establishing baseline behaviour based on observed behaviour for AI Agents might never be enough in case of AI Agents because of their autonomous nature and the ability to improve. The authors should look at other well-established techniques as well to see how they can use those to arrive at the baseline. For example, adversarial Red Teaming/stress testing of AI Agents can help you a lot to understand how an agent might behave in a certain scenario

    4) Another area where author/s should focus is the problem that enterprises face with UEBA deployments today and how to avoid them in case of AI Agents. For example, a high false positive rate and how the system will behave in case of false alarms

    Read full reviewShow less
  3. This submission rightfully highlights the need to not recreate security tooling from first principles, but rather adapt existing tooling with AI in mind. Using traditional security approaches, we can establish that behavioural profiles could work well for evaluation infrastructure, just as they have done for user and system behaviour.

    Agents can create various objects and take actions that may not look obviously malicious. A misaligned model may not simply write malicious code, but instead use something such as the Windows API in a way that resembles legitimate execution. As the project highlights, tracing the relationships between processes, files, network connections, system calls, and other resources could allow incident response teams to better understand the agent’s actions, identify suspicious behaviour, and potentially detect attempts at evasion. The individual action may look legitimate, but the sequence of actions and the relationships between them may tell a very different story.

    The paper is a great foundation to start from and I think it has the right ideas. I would like to see this put into practice to understand what works well, what doesn’t, and how it can be improved. There are a lot of directions this work could grow in, particularly once it is tested against real evaluation infrastructure and agent behaviour.

    Read full reviewShow less

Cite this project

@misc{sharma2026maintaining,
  title = {{MAINTAINING BEHAVIORAL PROFILES FOR ENTITIES IN AI EVALUATION INFRASTRUCTURE}},
  author = {Sandeep Sharma},
  year = {2026},
  month = sep,
  note = {Submitted to AI Incident Response Sprint, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/maintaining-behavioral-profiles-for-entities-in-ai-evaluation-infrastructure-f9wx}},
  url = {https://apartresearch.com/sprints/projects/maintaining-behavioral-profiles-for-entities-in-ai-evaluation-infrastructure-f9wx}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026