RAG Faithfulness evaluator
Mouad Benmansour · Team Faith Layer
Submitted to AI Control Hackathon 2026. Sprint projects are early-stage work by participants, not Apart Research publications.
RAG is everywhere now, and organisations adopt it precisely because it is supposed to be accurate. This tool holds RAG systems accountable to that promise, monitoring every output at sentence level, visualising how faithfulness drifts across an answer, and producing a named diagnosis explaining what went wrong, where, and what to fix. Built on a hybrid of semantic similarity and lexical overlap, with an LLM diagnostic layer constrained to a human-authored taxonomy of six known drift patterns. It catches what aggregate scores miss and explains what they cannot

Reviews
For a hackathon build, this is a clean working prototype. Sentence-level faithfulness scoring with drift timelines beats aggregate numbers — operators get something actionable, not just a pass/fail. The six-pattern diagnostic taxonomy is practical and the ACL injury catch (high semantic similarity but low lexical overlap revealing a training data leak) is a nice demo of why two scoring signals beat one.
The AI control connection could be tighter — this reads more as a RAG reliability tool than a control protocol for adversarial subversion. Framing it as "what if the model is deliberately trying to smuggle ungrounded claims past the monitor" would have landed better for this hackathon's theme.
Evaluation is one domain with hand-tuned thresholds, which is fine for a prototype — the roadmap (HaluEval, entailment scoring) shows the author knows what's next. Good execution for the time constraint.
Read full reviewShow less
The problem of sentence-level faithfulness scoring in RAG systems is practically relevant, and the drift timeline idea is an intuitive framing. The diagnostic taxonomy of six failure patterns is a reasonable starting point for categorising how RAG systems fail.
The main concern is fit with the AI control hackathon theme. This addresses ordinary retrieval tooling failures rather than the adversarial monitoring setting of AI control.
Cite this project
@misc{benmansour2026rag,
title = {{RAG Faithfulness evaluator}},
author = {Mouad Benmansour},
year = {2026},
month = mar,
note = {Submitted to AI Control Hackathon 2026, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/rag-faithfulness-evaluator-mnv0}},
url = {https://apartresearch.com/sprints/projects/rag-faithfulness-evaluator-mnv0}
}More from AI Control Hackathon 2026
- 1st placeLinuxArena track winnerView project: Omission Attacks: When Doing Nothing Is the Attack
Omission Attacks: When Doing Nothing Is the Attack
MAIA
AI control protocols monitor agent actions to detect sabotage, but omission attacks exploit what the agent fails to do rather than what it does. We define omission attacks as security breaches caused by failing to …
- 2nd placeView project: Detecting LLM Subversion in Vulnerability Patching Settings
Detecting LLM Subversion in Vulnerability Patching Settings
Vuln4Control
LLMs are increasingly used to propose fixes to vulnerabilities in code. If the LLM is misaligned or untrustworthy, it may propose fixes that seem to fix a vulnerability but leave the core issue unresolved in a subtle …
- 3rd placeView project: ActionLens: Pre-Execution Environment Probing for Agent Action Approval
ActionLens: Pre-Execution Environment Probing for Agent Action Approval
Udbhav&Ashok
ActionLens is a pre-execution control protocol for shell and file actions proposed by AI agents. Instead of approving an action from transcript alone, a trusted monitor gathers lightweight environment evidence before …