Encouraging Chain-of-Thought Reasoning
Shreyans Jain, Thomas Walker, Kutay Buyruk, Soumyadeep Bose · Team BlueDot Impact - Shreyans
Submitted to Reprogramming AI Models Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.
Encouraging Chain-of-Thought Reasoning via Feature Steering in Large Language Models
Reviews
Very cool stab at increasing CoT via steering. I would like to see a fuller investigation of how the faithfulness of the steered CoT compares to prompted CoT.
This is a really nice project on chain of throught. The experiments are logical and well conducted, and the presentation of the results is clear. The uplift in chain of thought performance is quite surprising - I'd be interested to know if the authors tuned the feature strengths or set them at the default intervention strength. Feature steering curves (feature strength vs performance) often peak at somewhat different points on different features (even semantically very similar ones) so tuning can be very worth doing. The findings on uncertainty at the first tokens of a direct response are intriguing and worth some more investigation.
A very interesting extension would be to test the generalisation of these features to another domain where CoT reasoning is important (ideally something non-mathematical, for example logic puzzles). Seeing a scatter plot of performance
on one domain vs performance on another domain would be very informative - my concern is that steering might improve one kind of performance at the expense of another.
Read full reviewShow less
cool idea with relevance to AI safety (model oversight / reasoning transparency; though slight caveat regarding faithfulness of CoT). I think this deserves further exploration and could potentially shed light on important methodological questions (such as faithfulness of model reasoning). These questions are not easy to study but this seems like a great first step!
Cite this project
@misc{jain2024encouraging,
title = {{Encouraging Chain-of-Thought Reasoning}},
author = {Shreyans Jain and Thomas Walker and Kutay Buyruk and Soumyadeep Bose},
year = {2024},
month = nov,
note = {Submitted to Reprogramming AI Models Hackathon, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/encouraging-chain-of-thought-reasoning}},
url = {https://apartresearch.com/sprints/projects/encouraging-chain-of-thought-reasoning}
}More from Reprogramming AI Models Hackathon
- 1st place by peer reviewView project: AutoSteer: Weight-Preserving Reinforcement Learning for Interpretable Model Control
AutoSteer: Weight-Preserving Reinforcement Learning for Interpretable Model Control
AI Safety Initiative Groningen
Traditional fine-tuning methods for language models, while effective, often disrupt internal model features that could provide valuable insights into model behavior. We present a novel approach combining Reinforcement …
- View project: Classification on Latent Feature Activation for Detecting Adversarial Prompt Vulnerabilities
Classification on Latent Feature Activation for Detecting Adversarial Prompt Vulnerabilities
Feature Disruption Lab
We present a method leveraging Sparse Autoencoder (SAE)-derived feature activations to identify and mitigate adversarial prompt hijacking in large language models (LLMs). By training a logistic regression classifier on …
- View project: Utilitarian Decision-Making in Models - Evaluation and Steering
Utilitarian Decision-Making in Models - Evaluation and Steering
Byte To Brain
We design an eval based on the Oxford Utilitarianism Scale (OUS) that measures the model’s deontological vs utilitarian preference in a nine question, two factor model. Using this scale, we measure how feature steering …