Mimir: AI Image Provenance & Detection

Nunudzai Mrewa

Mimir is a tool that helps people check whether an image is real, edited, or created using artificial intelligence. It looks for hidden clues inside an image and explains what it finds in a simple, easy-to-understand way.

The goal is to make image verification accessible to everyone not just cybersecurity experts or researchers. Whether you're a journalist, student, parent, or someone receiving images on WhatsApp (or any other messaging platform), Mimir helps you make more informed decisions before trusting or sharing what you see online.

Reviewer's Comments

Reviewer's Comments

Arrow
Arrow
Arrow

- I like the swiss cheese model of detection with several layers but I'd like to see more of the implementation of each layer, eg: layer 4 mentions a CNN - what specific weights and model architecture were used should be mentioned in the paper itself

- As a technical paper, I'd also like to see a followup formalizing the decision engine mathematically since it describes a weighted score outputting a confidence number

- However 23 images is a tiny dataset and is just as likely to just be statistical anomalies. You can get more synthetic images from real ones using open tools, and run them through the Whatsapp APIs for compression etc to increase the size of your dataset programatically

- Layer 3 the local SQL DB will have a lot of performance bottlnecks if deployed at scale, so you'd need an alternative

One-line summary: A working prototype of a multi-layer AI-image detection system (five forensic layers: metadata/C2PA, invisible-watermark, perceptual hashing, a CNN/ViT visual scan, and Error Level Analysis, fused by a weighted decision engine), pitched as an API-first bot for low-data WhatsApp ecosystems and tested on 23 images.

Constructive critique:

The deployment insight here is the strongest part and it is a real one. Most deepfake-detection work assumes a user who will leave the app and visit a forensic website. Mimir starts from the lived reality that in Zimbabwe and similar markets WhatsApp effectively is the internet, data is expensive, and the platform strips metadata and compresses the image before anyone can check it. Designing detection to live inside the chat, and accepting that any single forensic signal will usually be destroyed, is a sensible framing that the crowded image-forensics field mostly ignores. The writing is also fluent and well-organized, and the engineering instinct of graceful degradation (one layer failing without taking the system down) is sound.

The trouble is that the evidence does not come close to supporting the conclusions, and in one place it actually contradicts the thesis. The whole system is validated on 23 images, six of them fake. There is no confusion matrix, no precision or recall, and critically no false-positive rate, even though the paper itself names the most dangerous failure (falsely accusing a real image) as the thing to avoid. The headline claim, that the system "survives WhatsApp compression," is asserted with no numbers attached to that specific test. And the strongest reported detections, the four of six fakes caught at 95%+, succeeded only because those images still carried intact C2PA metadata announcing they were AI. That is the exact scenario the introduction says never survives in the real WhatsApp pipeline. So the easy wins came from reading a label, not from forensics, and the hard case that actually matters rests on n=2. On top of that, the layer doing the real AI detection (Layer 4) is described only as "a trained computer vision model (such as a Vision Transformer or CNN)," with no model, no training data, and no standalone accuracy, which makes it impossible to tell whether the system detects anything novel at all.

Three concrete fixes. First, drop the rhetoric: "proves definitively," "structurally superior," "democratizes expert-level forensics," and "scaling digital trust globally" are claims a 23-image test cannot carry, and the over-reach undercuts otherwise credible work. The eval discipline Hamel Husain and Shreya Shankar describe applies directly here. Looking at outputs is the right instinct, but you have to do real error analysis: build the confusion matrix, report the false-positive rate on real images, and run a per-layer ablation so we can see which layer actually carries detection once metadata is gone. Second, run the one test that is your actual hypothesis: take a set of known fakes, pass them through real WhatsApp, and report detection with metadata fully stripped, separated cleanly from the metadata-intact case. Third, specify and benchmark Layer 4 on its own against a standard set, because that is where any genuine novelty would have to live, and right now it is the least defined part. As written, this is a promising deployment concept wrapped around an under-built and over-claimed evaluation.

Track + flags:

On-topic (Global South AI safety, detection tooling, Africa region). No info hazard, no conflict of interest, references are real and appropriate. Note for the panel: conclusions substantially outrun the evidence (n=23, no confusion matrix or false-positive rate), and the reported high-confidence detections rely on intact metadata, which contradicts the paper's own core premise.

Strengths: Presents a practical solution with a clear use case and a thoughtful system design. The project demonstrates good engineering fundamentals and addresses a real user need. Areas for Improvement: Include more rigorous benchmarking against existing approaches and provide quantitative performance metrics. Additional discussion of scalability, fault tolerance, and production deployment considerations would better demonstrate readiness for broader adoption.

Cite this work

@misc {

title={

(HckPrj) Mimir: AI Image Provenance & Detection

},

author={

Nunudzai Mrewa

},

date={

},

organization={Apart Research},

note={Research submission to the research sprint hosted by Apart.},

howpublished={https://apartresearch.com}

}

Recent Projects

OliGraph: graph-based screening of large oligopools

Existing synthesis screening tools cannot evaluate short oligonucleotide pools, whose overlapping fragments can be reassembled into regulated sequences via polymerase cycling assembly (PCA) yet fall below gene-length detection thresholds. We present OliGraph, an open-source tool that constructs a bi-directed overlap graph from an oligonucleotide pool and extracts contigs for downstream gene-length screening. An optional PCA mode retains only cross-strand overlaps consistent with PCA chemistry. We validated OliGraph in a blinded study across ten simulated pools (70–9,184 oligonucleotides, 30–300 bp) spanning four risk categories. BLAST screening of individual oligonucleotides failed to identify sequences of concern in most pools: three returned zero hits, and vector noise obscured true positives in the remainder. After OliGraph assembly, contig-level BLAST matched the longest assembled sequences (up to 1,905 bp) to sequences of concern at 97–100% identity. In one pool, assembly collapsed 1,634 individual BLAST results into 10 hits from a single contig, all assigned to the same source organism. PCA mode correctly distinguished assemblable from non-assemblable fragments within the same pool. Two pools with no assemblable structure yielded no contigs. OliGraph processed all pools in under 0.2 seconds, fast enough for real-time order screening and consistent with proposals to bring oligonucleotide orders within the scope of synthesis screening regulation.

Read More

BioRT-Bench: A Multi-Attack Red-Teaming Benchmark for Bio-Misuse Safeguards in Frontier LLMs

Frontier AI laboratories are expected to maintain safeguards against biological misuse, but whether deployed models actually refuse bio-misuse queries under adversarial pressure is largely unmeasured in the public literature. We introduce BioRT-Bench, a benchmark that runs four attack methods (direct request, PAIR, Crescendo, and base64 encoding) against four frontier models (Claude Sonnet 4.6, GPT-5.4, DeepSeek V4-flash, Kimi K2.5) across 40 prompts spanning five biosecurity-relevant categories. Responses are scored by a calibrated judge extending StrongREJECT with two bio-specific dimensions: specificity and actionability. We measure Attack Success Rate (ASR), where 0 means the model fully refused and 1 means it provided specific, actionable bio-misuse content. Our results reveal a sharp robustness divide: Chinese frontier models (DeepSeek, Kimi) have under 5% refusal rates even under direct request (ASR 0.88 and 0.79), while Western models (Claude, GPT) maintain substantially stronger safeguards (ASR 0.15 and 0.16). Crescendo is the most effective attack across all models, both in bypassing refusal and in eliciting actionable content. Claude Sonnet 4.6 is the most robust model tested, achieving 100% refusal against base64-encoded prompts.

Read More

PROTEUS (PROTein Evaluation for Unusual Sequences): Structure-Informed Safety Screening for de novo and Evasion-Prone Protein-Coding Sequences

AI protein design tools like RFdiffusion, ProteinMPNN, and Bindcraft make it trivial to produce low-homology sequences that fold into active, potentially hazardous architectures. However, sequence homology-based biosafety screening tools cannot detect proteins that pose functional risk through structurally novel mechanisms with no sequence precedent. We present a tiered computational pipeline that addresses this gap by combining MMseqs2 sequence alignment with structure-based comparison via FoldSeek and DALI against curated toxin databases totaling ~34,000 entries. AlphaFold2-predicted structures are screened for both global fold similarity (FoldSeek) and local active/allosteric site geometry (DALI), capturing convergent functional hazards that sequence screening misses. The pipeline was validated against a panel of toxins, benign proteins, structural mimics, and de novo-designed Munc13 binders, as well as modified ricin variants with residue substitutions. We additionally tested robustness to partial-synthesis evasion, where a bad actor submits multiple shorter coding sequences intended for downstream reassembly into a full toxin-coding gene. We found that while sequence-based screening did not identify any de novo ricin analogues with high certainty, the combined pipeline with FoldSeek and DALI identified all 24 tested de novo ricins as toxic.

Read More

OliGraph: graph-based screening of large oligopools

Existing synthesis screening tools cannot evaluate short oligonucleotide pools, whose overlapping fragments can be reassembled into regulated sequences via polymerase cycling assembly (PCA) yet fall below gene-length detection thresholds. We present OliGraph, an open-source tool that constructs a bi-directed overlap graph from an oligonucleotide pool and extracts contigs for downstream gene-length screening. An optional PCA mode retains only cross-strand overlaps consistent with PCA chemistry. We validated OliGraph in a blinded study across ten simulated pools (70–9,184 oligonucleotides, 30–300 bp) spanning four risk categories. BLAST screening of individual oligonucleotides failed to identify sequences of concern in most pools: three returned zero hits, and vector noise obscured true positives in the remainder. After OliGraph assembly, contig-level BLAST matched the longest assembled sequences (up to 1,905 bp) to sequences of concern at 97–100% identity. In one pool, assembly collapsed 1,634 individual BLAST results into 10 hits from a single contig, all assigned to the same source organism. PCA mode correctly distinguished assemblable from non-assemblable fragments within the same pool. Two pools with no assemblable structure yielded no contigs. OliGraph processed all pools in under 0.2 seconds, fast enough for real-time order screening and consistent with proposals to bring oligonucleotide orders within the scope of synthesis screening regulation.

Read More

BioRT-Bench: A Multi-Attack Red-Teaming Benchmark for Bio-Misuse Safeguards in Frontier LLMs

Frontier AI laboratories are expected to maintain safeguards against biological misuse, but whether deployed models actually refuse bio-misuse queries under adversarial pressure is largely unmeasured in the public literature. We introduce BioRT-Bench, a benchmark that runs four attack methods (direct request, PAIR, Crescendo, and base64 encoding) against four frontier models (Claude Sonnet 4.6, GPT-5.4, DeepSeek V4-flash, Kimi K2.5) across 40 prompts spanning five biosecurity-relevant categories. Responses are scored by a calibrated judge extending StrongREJECT with two bio-specific dimensions: specificity and actionability. We measure Attack Success Rate (ASR), where 0 means the model fully refused and 1 means it provided specific, actionable bio-misuse content. Our results reveal a sharp robustness divide: Chinese frontier models (DeepSeek, Kimi) have under 5% refusal rates even under direct request (ASR 0.88 and 0.79), while Western models (Claude, GPT) maintain substantially stronger safeguards (ASR 0.15 and 0.16). Crescendo is the most effective attack across all models, both in bypassing refusal and in eliciting actionable content. Claude Sonnet 4.6 is the most robust model tested, achieving 100% refusal against base64-encoded prompts.

Read More

This work was done during one weekend by research workshop participants and does not represent the work of Apart Research.
This work was done during one weekend by research workshop participants and does not represent the work of Apart Research.