Skip to content
Sprint projectFeb 1, 2026Tirana

ATrain

ATrain · Team Ajsel Budlla

Submitted to The Technical AI Governance Challenge. Sprint projects are early-stage work by participants, not Apart Research publications.

ATrain is a system that lets us prove, in a secure and transparent way, how much compute a model actually used during training. It logs training metrics, estimates computational usage, and cryptographically signs the data so anyone can verify it hasn’t been tampered with. We built a simple dashboard in Colab where you can start a training run and generate graphs showing whether the model stayed within regulatory thresholds. Essentially, it turns AI training from a trust-based process into one that can be independently verified without exposing your model weights or sensitive data. This could help labs and regulators ensure safe scaling of AI systems.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for AI safety if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical AI safety problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. The paper correctly identifies that current approaches in compute governance rely on voluntary disclosure rather than verifiable evidence, and helps solve one part of the problem of verification. The emphasis on an accessible Colab-based dashboard is a good design choice for a prototype, lowering the barrier for non-technical stakeholders to interact with attestation concepts, which has value for building intuition about what compute governance tooling could look like. However, the system is at present self-attesting: the training process reports its own compute, then cryptographically signs its own report. The signature proves logs weren't altered after signing, but says nothing about whether they were accurate in the first place, with a dishonest actor able to produce false logs and sign those. The paper claims to transform training "from trust-based to proof-based," but it relocates the trust assumption up the chain rather than eliminating it. This is a useful first layer in what would need to be a multi-layered verification stack, but other components, such as a TEE or hardware-level measurement providing an external root of trust, would still be necessary for cryptographic reporting of compute usage to be reliable.

    Read full reviewShow less
  2. This is well built and easy to understand. The dashboard + workflow is clean.

    But the core idea (hash + sign logs to prove integrity) is a pretty standard pattern. So the novelty is limited. Also: it mainly proves “the log wasn’t changed,” not “the training run was honest.” If the environment lies, a signed log can still be fake.

    To make this more impactful, I want a clear threat model and a real path to hardware-backed / remote attestation. Right now it’s good compliance tooling, but not a new safety mechanism.

Cite this project

@misc{atrain2026atrain,
  title = {{ATrain}},
  author = {ATrain},
  year = {2026},
  month = feb,
  note = {Submitted to The Technical AI Governance Challenge, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/atrain-mvbo}},
  url = {https://apartresearch.com/sprints/projects/atrain-mvbo}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026