Skip to content
Sprint projectFeb 2, 2026Chennai, India

Attested Multi-Agent Conversation Logs: A Tamper-Evident Black Box for AI Governance

Anantha Shakthi Ganeshan Thevar, Publius Dirac · Team Attested Logs

Submitted to The Technical AI Governance Challenge. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: Attested Multi-Agent Conversation Logs: A Tamper-Evident Black Box for AI Governance

Code (opens in new tab)
Share

As multi-agent AI systems assume high-stakes responsibilities, governance frameworks such as the EU AI Act demand traceable, auditable records of agent interactions. We introduce Attested Logs, an open-source Python library that serves as a tamper-evident "black box" for AI conversations. Each message is cryptographically signed (Ed25519), hash-chained (SHA-256), and anchored to trusted public keys, enabling fully offline verification of integrity, authenticity, and ordering. Inspired by aviation flight recorders and C2PA content provenance, the library supports incident forensics, regulatory compliance, and cross-organizational trust. Integrations for AutoGen (manual logging) and LangGraph (callback-based auditing) are provided, together with runnable demos demonstrating end-to-end signing, verification, and tamper detection in real LLM conversations, and 20+ passing tests covering core, crypto, chain, and verification layers.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for AI safety if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical AI safety problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. This is the most shippable project here. It actually delivers a real library that people could use tomorrow: signed messages, hash chaining, offline verification, integrations, tests. For a hackathon, that’s exceptional execution.

    Novelty is moderate because the crypto pieces are standard. The value is that you packaged it cleanly and made it usable in real agent frameworks. This is exactly the kind of infrastructure that improves auditability in deployed systems.

    To go further: key management (rotation / revocation), what happens when an agent is compromised, and scaling / storage patterns. But as a sprint project: this is very strong.

  2. Tamper-proof evidence of multi-agent interactions is important for future digital forensics and AI governance and regulation. So the project definitely explores an important problem. The write-up is clear on the purely technical side as well as on some of the related work. And the code artifact seems solid.

    To make this a much stronger contribution: Explore one clearly defined actual threat model in depth. Red-team your approach and figure out what the hard part actually is. More in-depth exploration would probably reveal that the trust chain bottoming out at the model provider is a fundamental limitation of the approach your describing. Put yourself in the shoes of an adversary (eg a model provider who wants to alter the logs): What attacks can they run and how does your approach fare?

Cite this project

@misc{thevar2026attested,
  title = {{Attested Multi-Agent Conversation Logs: A Tamper-Evident Black Box for AI Governance}},
  author = {Anantha Shakthi Ganeshan Thevar and Publius Dirac},
  year = {2026},
  month = feb,
  note = {Submitted to The Technical AI Governance Challenge, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/attested-multiagent-conversation-logs-a-tamperevident-black-box-for-ai-governance-du9i}},
  url = {https://apartresearch.com/sprints/projects/attested-multiagent-conversation-logs-a-tamperevident-black-box-for-ai-governance-du9i}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026