Skip to content
Sprint projectNov 24, 2024

Encouraging Chain-of-Thought Reasoning

Shreyans Jain, Thomas Walker, Kutay Buyruk, Soumyadeep Bose · Team BlueDot Impact - Shreyans

Submitted to Reprogramming AI Models Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: Encouraging Chain-of-Thought Reasoning

Code (opens in new tab)
Share

Encouraging Chain-of-Thought Reasoning via Feature Steering in Large Language Models

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

Does the project contribute to the field of mechanistic interpretability? Does it provide new insights into understanding or steering AI model behavior? How well does it move us towards reprogramming AI models? How original and innovative is the approach?

How important is the contribution to advancing the field of AI safety? Do we expect the results to generalize beyond the specific case(s) presented in the submission? Does the approach introduce new safety mechanisms or enhance existing ones in innovative ways?

How well is the project executed from a technical standpoint?, Is the code well-structured, documented, and reproducible?, How effectively does it utilize Goodfire's SDK/API and other provided resources?, How clearly and effectively is the research presented in the paper?, Quality of visualizations and demos (if applicable), Clarity of methodology explanation and results interpretation

  1. Very cool stab at increasing CoT via steering. I would like to see a fuller investigation of how the faithfulness of the steered CoT compares to prompted CoT.

  2. This is a really nice project on chain of throught. The experiments are logical and well conducted, and the presentation of the results is clear. The uplift in chain of thought performance is quite surprising - I'd be interested to know if the authors tuned the feature strengths or set them at the default intervention strength. Feature steering curves (feature strength vs performance) often peak at somewhat different points on different features (even semantically very similar ones) so tuning can be very worth doing. The findings on uncertainty at the first tokens of a direct response are intriguing and worth some more investigation.

    A very interesting extension would be to test the generalisation of these features to another domain where CoT reasoning is important (ideally something non-mathematical, for example logic puzzles). Seeing a scatter plot of performance

    on one domain vs performance on another domain would be very informative - my concern is that steering might improve one kind of performance at the expense of another.

    Read full reviewShow less
  3. cool idea with relevance to AI safety (model oversight / reasoning transparency; though slight caveat regarding faithfulness of CoT). I think this deserves further exploration and could potentially shed light on important methodological questions (such as faithfulness of model reasoning). These questions are not easy to study but this seems like a great first step!

Cite this project

@misc{jain2024encouraging,
  title = {{Encouraging Chain-of-Thought Reasoning}},
  author = {Shreyans Jain and Thomas Walker and Kutay Buyruk and Soumyadeep Bose},
  year = {2024},
  month = nov,
  note = {Submitted to Reprogramming AI Models Hackathon, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/encouraging-chain-of-thought-reasoning}},
  url = {https://apartresearch.com/sprints/projects/encouraging-chain-of-thought-reasoning}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026