TensorGuard-Lite: Auditing Sovereign AI Claims Through Gradient-Based Model Provenance
Adarsh Mishra · Team Abhinav
Submitted to Global South AI Safety Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.
# TensorGuard-Lite
## Problem
* No reliable way to verify if "Sovereign AI" models are genuinely indigenous or fine-tuned foreign models ("open-washing"). * Lack of accountability in state-funded AI compute programs.
## Solution
* White-box AI provenance auditor for open-weight LLMs. * Uses gradient fingerprints, tokenizer analysis, and governance scoring. * Runs on Google Colab T4 (16GB VRAM).
## Core Method
### 1. Gradient Fingerprinting
* Extracts deterministic 16-dimensional fingerprints. * Uses attention, FFN, embedding, and structural features. * Compares models using cosine similarity and Euclidean distance.
### 2. Token Tax Analysis
* Measures tokenizer fertility across:
* English * Hindi * Tamil * Estimates token inflation and attention-cost overhead.
### 3. Governance Scorecard
* 10 transparency dimensions. * Generates risk level:
* Low * Medium * High * Unverifiable
## Models Audited
* Llama-3.2-1B * SmolLM2-1.7B-Instruct * Qwen2.5-1.5B * DeepSeek-R1-Distill-Qwen-1.5B
## Key Features
* Training-free auditing * Deterministic fingerprints * PCA lineage clustering * Cosine + Euclidean similarity * Tokenizer Jaccard similarity * Governance transparency scoring * Interactive Gradio dashboard * JSON / CSV / LaTeX / PNG / SVG exports
## Key Findings
* DeepSeek clusters near Qwen, supporting lineage detection. * API-only models remain difficult to verify. * Indic languages show significantly higher tokenization costs than English.
## Limitations
* Requires safetensors / white-box access. * Similar architectures may create false positives. * Small tokenizer corpus. * Best suited for models ≤3B parameters.
## Policy Proposal
**Compute-Conditional Disclosure Policy**
Organizations receiving public AI compute subsidies should disclose:
* Data lineage proofs * Tokenizer vocabulary * Model weights for regulators * Architecture documentation * Safety evaluation results
## Implementation
* `tensorguard_lite_colab.py` * `TensorGuard_Lite_Colab.ipynb` * Google Colab (T4 GPU) * GPL v3 License
## Repository
https://github.com/Adarsh-Me/Sovergien
Reviews
The idea is pretty original! provenance fingerprinting as an enforcement lever tied to state compute. And you ran it which shows well. Recovering the known DeepSeek/Qwen lineage on a T4 (cosine 0.980, vocab Jaccard 0.9999) is a real, reproducible result. Where it thins out is validation. Four models and one known derivative is not enough to bound false positives, and your own numbers prove the risk: SmolLM2's cosine to Qwen (0.997) beats the actual derivative's (0.980), so ranking on cosine alone mislabels it. Widen the panel, add several known derivatives, report false-positive rates. Lead with the cleaner Jaccard signal and treat gradient cosine as backup. And since white-box access is the thing sovereign vendors will not give you, the compute-conditional disclosure policy is your real contribution, so build that out. Promising, and the policy framing is the part to chase.
Creative application of gradient fingerprinting. The Tamil Token Tax finding is striking enough that it deserves more than a subsection. I think it's arguably the most immediately actionable finding for policymakers deciding whether to mandate indigenous tokenizers as a condition of compute subsidy. You may consider addressing the paper's fundamental tension more directly: that gradient fingerprinting requires white-box weight access, but the models most urgently needing audit are precisely those refusing to release weights. Could be by proposing either black-box behavioral fingerprinting alternatives or specific legal mechanisms that would compel regulator-level weight access as a condition of receiving state compute.
Adapting gradient-based fingerprinting to the open-washing problem is a sharp framing, and the system behind it is real: a deterministic 2,097-line Colab pipeline, a 16-dimensional fingerprint, and an unusually clean two-signal result, where DeepSeek-R1-Distill-Qwen lands at cosine 0.980 to its Qwen parent and a near-perfect 0.9999 tokenizer-vocabulary Jaccard corroborates the lineage independently. The limitations are honest, naming false positives on similar architectures, the twelve-sentence corpus, and the ≤3B ceiling. Validation is the soft spot. The one positive pair is a known distill, and the threshold (cosine ≥0.92) is set to pass exactly that pair, so the discriminator rests on n=1; SmolLM2's 0.997 cosine to Qwen — higher than the true derivative's — is a near-false-positive that only Euclidean distance rescues, which is fragile. More fundamentally, the method needs white-box weights, yet the sovereign models it targets are API-only and self-classify as UNVERIFIABLE, so the tool works where it is least needed and cannot reach where it is most needed. An external set of genuinely open-washed models, with a measured false-positive rate across unrelated same-architecture pairs, is what would make this a detector rather than a demonstration.
Read full reviewShow less
Cite this project
@misc{mishra2026tensorguardlite,
title = {{TensorGuard-Lite: Auditing Sovereign AI Claims Through Gradient-Based Model Provenance}},
author = {Adarsh Mishra},
year = {2026},
month = jun,
note = {Submitted to Global South AI Safety Hackathon, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/tensorguardlite-auditing-sovereign-ai-claims-through-gradientbased-model-provenance-23yh}},
url = {https://apartresearch.com/sprints/projects/tensorguardlite-auditing-sovereign-ai-claims-through-gradientbased-model-provenance-23yh}
}More from Global South AI Safety Hackathon
- View project: Pragmatic Sophistry in Vietnamese Multi-Agent Oversight
Pragmatic Sophistry in Vietnamese Multi-Agent Oversight
AI Safety Enthusiasts
AI safety monitors are usually evaluated on the assumption that risky behavior is lexically visible in the text being watched. We test this assumption in a multilingual, multi-agent setting: Vietnamese-language workflow …
- View project: JusticIA: A Counterfactual Benchmark for Auditing Contextual Biases in Language Models for Transitional Justice
JusticIA: A Counterfactual Benchmark for Auditing Contextual Biases in Language Models for Transitional Justice
JusticeMiners
JusticIA is a counterfactual benchmark for auditing contextual bias in LLMs applied to Colombian transitional justice. It tests whether six LLMs change their sanction recommendations when only one contextual attribute …
- View project: Coldron
Coldron
ColDron
En Colombia, los grupos armados ilegales ya atacan con drones comerciales modificados y ya han herido y matado a civiles. Una pregunta decide cómo gobernar esta amenaza: ¿quién elige el blanco y aprieta el gatillo? Hoy, …