Eidolon
Mwelwa Kashingwa
This project introduces **EIDOLON** as a framework for studying **AI identity continuity** under component replacement. Three models (Llama, Gemma, and Qwen) were tested with varying levels of **model replacement** and **memory replacement** ranging from 0% to 100%. Identity was evaluated using five consistency questions, producing an **Identity Score** that was averaged into an **Identity Continuity Score (ICS)**. Results showed that continuity is not tied to a single component but emerges from the interaction between memory and model weights, with surprisingly high values even under heavy replacement. These findings highlight both the resilience of identity signals in AI systems and the methodological limits of current evaluation approaches, with implications for **AI safety** and system updates.
Overall good preliminary exploration and framework for future investigations. The formulaic methods are noteworthy. The idea of replacing parts, and the connection to the "Ship of Theseus" metaphor is a ripe comparision which is generally well situated within the emerging field. Seed numbers selection is reasonable but statistical rigor could be introduced. Methods could expand into other replacement pieces beyond the preliminary parts used in this paper, which are appropriate for a weekend sprint. The writing was difficult to read with odd colon statements, in-line emphasis, and artificial phrasing which don't match an academic paper submission. I would recommend edits including rewording overly artificial sounding language, featuring the ship of theseus metaphor more prominently throughout, noting the key findings, general methods, and highlight of result in the abstract, as well as referencing citations within the paper more broadly.
Ferreira and Kashingwa play a bit of a game with Qwen 2.5 3B, Llama 3.2 3B, and Gemma 3 4B. They tell each model about a hypothetical EIDOLON_0 AI whose model name and memories are gradually replaced. The point is to find out what the models say about the Ship of Theseus as applied to AI especially when the EIDOLON_0 is changed to qwen2.5:3b, llama3.2:3b, gemma3:4b. Alas those little models didn't pick up on the joke. Across 680 responses, the authors' data show that when asked directly the models say the underlying model is important for identity. But when asked to rate changes, memory-swaps move the needle and model swaps don't. Probably can't expect much subtlety from such small models.
Cite this work
@misc {
title={
(HckPrj) Eidolon
},
author={
Mwelwa Kashingwa
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


