Skip to content
Sprint projectJan 24, 2025

CoTEP: A Multi-Modal Chain of Thought Evaluation Platform for the Next Generation of SOTA AI Models

Alyssia J, Martin CL · Team CoTEP

Submitted to AI Safety Entrepreneurship Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: CoTEP: A Multi-Modal Chain of Thought Evaluation Platform for the Next Generation of SOTA AI Models

More on youtu.be (opens in new tab)
Share

As advanced state-of-the-art models like OpenAI's o-1 series, the upcoming o-3 family, Gemini 2.0 Flash Thinking and DeepSeek display increasingly sophisticated chain-of-thought (CoT) capabilities, our safety evaluations have not yet caught up. We propose building a platform that allows us to gather systematic evaluations of AI reasoning processes to create comprehensive safety benchmarks. Our Chain of Thought Evaluation Platform (CoTEP) will help establish standards for assessing AI reasoning and ensure development of more robust, trustworthy AI systems through industry and government collaboration.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

  1. This is a super cool project! Really interesting to get expert-driven CoTs in for evaluation. There's a few questions regarding the impact on AI safety since it's a capability evaluation and will help to get stronger training data but the actual outlined strategy seems very reasonable. I highly suggest moving forward with this work and getting experimental data about existing CoT models, especially DeepSeek's R1 since it represents the next paradigm and CoT is visible. Great work.

  2. High-quality evaluations of chains of thought are an interesting opportunity. I'd love to see experiments in this direction!

Cite this project

@misc{j2025cotep,
  title = {{CoTEP: A Multi-Modal Chain of Thought Evaluation Platform for the Next Generation of SOTA AI Models}},
  author = {Alyssia J and Martin CL},
  year = {2025},
  month = jan,
  note = {Submitted to AI Safety Entrepreneurship Hackathon, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/cotep-a-multi-modal-chain-of-thought-evaluation-platform-for-the-next-generation-of-sota-ai-models}},
  url = {https://apartresearch.com/sprints/projects/cotep-a-multi-modal-chain-of-thought-evaluation-platform-for-the-next-generation-of-sota-ai-models}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026