LittleBrotherAI
Tarik Rosin, Rafi Hakim, Teodora Kamova, Manon Kempermann, Ruslan Dontsov · Team LittleBrotherAI
Submitted to Defensive Acceleration Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.
Large-language models (LLMs) enable powerful chatbots capable of chain-of-thought (CoT) reasoning and detailed responses. However, their reasoning often lacks transparency, making it difficult for users to judge reliability, detect hallucinations, or identify adversarial vulnerabilities [2,3,4]. In this work, we present a web-based chatbot platform “LittleBrotherAI” that integrates an external model with a comprehensive model evaluation dashboard. For every interaction, the system automatically assesses the model’s reasoning and response across seven key dimensions: consistency (language, semantics, and logical inference), clarity, reproducibility, legibility, coverage, adversarial behaviour, and accuracy. These evaluations are performed using dedicated auxiliary LLMs (Little brothers monitoring the main model) or specialised AI components[1]. The resulting scores are aggregated to determine a trustworthiness verdict for each answer, providing users with a clearer basis for judgment and trust calibration. Our platform demonstrates how transparent metrics and structured evaluation could enhance the responsible use of LLM-generated content. Our code can be found here: https://github.com/orgs/LittleBrotherAI/repositories.
Reviews
No public critique yet.
Cite this project
@misc{rosin2025littlebrotherai,
title = {{LittleBrotherAI}},
author = {Tarik Rosin and Rafi Hakim and Teodora Kamova and Manon Kempermann and Ruslan Dontsov},
year = {2025},
month = nov,
note = {Submitted to Defensive Acceleration Hackathon, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/littlebrotherai-yzzx}},
url = {https://apartresearch.com/sprints/projects/littlebrotherai-yzzx}
}More from Defensive Acceleration Hackathon
- View project: Neops - DevSecOps for the AI era
Neops - DevSecOps for the AI era
Broad Bros
NEOps is a CLI-based tool that embeds AI safety into your product lifecycle from day one. While development teams routinely build cybersecurity checks, AI-safety often comes later—or not at all. NEOps fills that gap by …
- View project: Assisted Audit of Solana Programs
Assisted Audit of Solana Programs
GLAM
Multi-agent solution that assists in auditing Solana programs, allows to consolidate audit findings into a knowledge base, and can integrate into CI/CD pipelines to prevent security regressions.
- View project: Mechanistic Watchdog
Mechanistic Watchdog
SL5
Mechanistic Watchdog is a mechanistic-interpretability-based “cognitive kill switch” for language models. Instead of only filtering final text, we monitor a model’s internal activations in real time and learn linear …