Skip to content
Sprint projectJun 22, 2026New Delhi, India

TensorGuard-Lite: Auditing Sovereign AI Claims Through Gradient-Based Model Provenance

Adarsh Mishra · Team Abhinav

Submitted to Global South AI Safety Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: TensorGuard-Lite: Auditing Sovereign AI Claims Through Gradient-Based Model Provenance

Code (opens in new tab)
Share

# TensorGuard-Lite

## Problem

* No reliable way to verify if "Sovereign AI" models are genuinely indigenous or fine-tuned foreign models ("open-washing"). * Lack of accountability in state-funded AI compute programs.

## Solution

* White-box AI provenance auditor for open-weight LLMs. * Uses gradient fingerprints, tokenizer analysis, and governance scoring. * Runs on Google Colab T4 (16GB VRAM).

## Core Method

### 1. Gradient Fingerprinting

* Extracts deterministic 16-dimensional fingerprints. * Uses attention, FFN, embedding, and structural features. * Compares models using cosine similarity and Euclidean distance.

### 2. Token Tax Analysis

* Measures tokenizer fertility across:

* English * Hindi * Tamil * Estimates token inflation and attention-cost overhead.

### 3. Governance Scorecard

* 10 transparency dimensions. * Generates risk level:

* Low * Medium * High * Unverifiable

## Models Audited

* Llama-3.2-1B * SmolLM2-1.7B-Instruct * Qwen2.5-1.5B * DeepSeek-R1-Distill-Qwen-1.5B

## Key Features

* Training-free auditing * Deterministic fingerprints * PCA lineage clustering * Cosine + Euclidean similarity * Tokenizer Jaccard similarity * Governance transparency scoring * Interactive Gradio dashboard * JSON / CSV / LaTeX / PNG / SVG exports

## Key Findings

* DeepSeek clusters near Qwen, supporting lineage detection. * API-only models remain difficult to verify. * Indic languages show significantly higher tokenization costs than English.

## Limitations

* Requires safetensors / white-box access. * Similar architectures may create false positives. * Small tokenizer corpus. * Best suited for models ≤3B parameters.

## Policy Proposal

**Compute-Conditional Disclosure Policy**

Organizations receiving public AI compute subsidies should disclose:

* Data lineage proofs * Tokenizer vocabulary * Model weights for regulators * Architecture documentation * Safety evaluation results

## Implementation

* `tensorguard_lite_colab.py` * `TensorGuard_Lite_Colab.ipynb` * Google Colab (T4 GPU) * GPL v3 License

## Repository

https://github.com/Adarsh-Me/Sovergien

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for AI safety if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical AI safety problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. The idea is pretty original! provenance fingerprinting as an enforcement lever tied to state compute. And you ran it which shows well. Recovering the known DeepSeek/Qwen lineage on a T4 (cosine 0.980, vocab Jaccard 0.9999) is a real, reproducible result. Where it thins out is validation. Four models and one known derivative is not enough to bound false positives, and your own numbers prove the risk: SmolLM2's cosine to Qwen (0.997) beats the actual derivative's (0.980), so ranking on cosine alone mislabels it. Widen the panel, add several known derivatives, report false-positive rates. Lead with the cleaner Jaccard signal and treat gradient cosine as backup. And since white-box access is the thing sovereign vendors will not give you, the compute-conditional disclosure policy is your real contribution, so build that out. Promising, and the policy framing is the part to chase.

  2. Creative application of gradient fingerprinting. The Tamil Token Tax finding is striking enough that it deserves more than a subsection. I think it's arguably the most immediately actionable finding for policymakers deciding whether to mandate indigenous tokenizers as a condition of compute subsidy. You may consider addressing the paper's fundamental tension more directly: that gradient fingerprinting requires white-box weight access, but the models most urgently needing audit are precisely those refusing to release weights. Could be by proposing either black-box behavioral fingerprinting alternatives or specific legal mechanisms that would compel regulator-level weight access as a condition of receiving state compute.

  3. Adapting gradient-based fingerprinting to the open-washing problem is a sharp framing, and the system behind it is real: a deterministic 2,097-line Colab pipeline, a 16-dimensional fingerprint, and an unusually clean two-signal result, where DeepSeek-R1-Distill-Qwen lands at cosine 0.980 to its Qwen parent and a near-perfect 0.9999 tokenizer-vocabulary Jaccard corroborates the lineage independently. The limitations are honest, naming false positives on similar architectures, the twelve-sentence corpus, and the ≤3B ceiling. Validation is the soft spot. The one positive pair is a known distill, and the threshold (cosine ≥0.92) is set to pass exactly that pair, so the discriminator rests on n=1; SmolLM2's 0.997 cosine to Qwen — higher than the true derivative's — is a near-false-positive that only Euclidean distance rescues, which is fragile. More fundamentally, the method needs white-box weights, yet the sovereign models it targets are API-only and self-classify as UNVERIFIABLE, so the tool works where it is least needed and cannot reach where it is most needed. An external set of genuinely open-washed models, with a measured false-positive rate across unrelated same-architecture pairs, is what would make this a detector rather than a demonstration.

    Read full reviewShow less

Cite this project

@misc{mishra2026tensorguardlite,
  title = {{TensorGuard-Lite: Auditing Sovereign AI Claims Through Gradient-Based Model Provenance}},
  author = {Adarsh Mishra},
  year = {2026},
  month = jun,
  note = {Submitted to Global South AI Safety Hackathon, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/tensorguardlite-auditing-sovereign-ai-claims-through-gradientbased-model-provenance-23yh}},
  url = {https://apartresearch.com/sprints/projects/tensorguardlite-auditing-sovereign-ai-claims-through-gradientbased-model-provenance-23yh}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026