CoTEP: A Multi-Modal Chain of Thought Evaluation Platform for the Next Generation of SOTA AI Models
Alyssia J, Martin CL · Team CoTEP
Submitted to AI Safety Entrepreneurship Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.
As advanced state-of-the-art models like OpenAI's o-1 series, the upcoming o-3 family, Gemini 2.0 Flash Thinking and DeepSeek display increasingly sophisticated chain-of-thought (CoT) capabilities, our safety evaluations have not yet caught up. We propose building a platform that allows us to gather systematic evaluations of AI reasoning processes to create comprehensive safety benchmarks. Our Chain of Thought Evaluation Platform (CoTEP) will help establish standards for assessing AI reasoning and ensure development of more robust, trustworthy AI systems through industry and government collaboration.
Reviews
This is a super cool project! Really interesting to get expert-driven CoTs in for evaluation. There's a few questions regarding the impact on AI safety since it's a capability evaluation and will help to get stronger training data but the actual outlined strategy seems very reasonable. I highly suggest moving forward with this work and getting experimental data about existing CoT models, especially DeepSeek's R1 since it represents the next paradigm and CoT is visible. Great work.
High-quality evaluations of chains of thought are an interesting opportunity. I'd love to see experiments in this direction!
Cite this project
@misc{j2025cotep,
title = {{CoTEP: A Multi-Modal Chain of Thought Evaluation Platform for the Next Generation of SOTA AI Models}},
author = {Alyssia J and Martin CL},
year = {2025},
month = jan,
note = {Submitted to AI Safety Entrepreneurship Hackathon, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/cotep-a-multi-modal-chain-of-thought-evaluation-platform-for-the-next-generation-of-sota-ai-models}},
url = {https://apartresearch.com/sprints/projects/cotep-a-multi-modal-chain-of-thought-evaluation-platform-for-the-next-generation-of-sota-ai-models}
}More from AI Safety Entrepreneurship Hackathon
- 1st place by peer reviewView project: AntiMidas: Building Commercially-Viable Agents for Alignment Dataset Generation
AntiMidas: Building Commercially-Viable Agents for Alignment Dataset Generation
the commonwealth
AI alignment lacks high-quality, real-world preference data needed to align agentic superintel- ligent systems. Our technical innovation builds on Pacchiardi et al. (2023)’s breakthrough in detecting AI deception …
- View project: Scoped LLM: Enhancing Adversarial Robustness and Security Through Targeted Model Scoping
Scoped LLM: Enhancing Adversarial Robustness and Security Through Targeted Model Scoping
FocusAI
Even with Reinforcement Learning from Human or AI Feedback (RLHF/RLAIF) to avoid harmful outputs, fine-tuned Large Language Models (LLMs) often present insufficient refusals due to adversarial attacks causing them to …
- View project: Prompt+question Shield
Prompt+question Shield
Seon's team
A protective layer using prompt injections and difficult questions to guard comment sections from AI-driven spam.