Let LLM Agents Perform LLM Surgery
Sharat Jacob Jacob · Team Agentic Kittens
Submitted to Reprogramming AI Models Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.
This project aimed to create and utilize LLM agents that could perform various mechanistic interventions on other LLMs. A few experiments were conducted ranging from an agent unsteering a mechanistically steered model to a neutral state, to an agent performing mechanistic edits to create a custom LLM as per user requirement. Goodfire's API was utilized along with their pre-defined functions to create the actions the agents would utilize.
Reviews
This is a cool idea - getting an agent to edit the internals of another agent to perform the task. As the authors observe, this is pretty difficult (even for frontier models) but as far as I can tell they had some successes. Quantifying these successes/failures can be difficult, especially at a relatively small scale and with time constraints - I'd encourage the authors to report a few transcripts of success & failure so we can get a feel for how well this performs and where steering agents succeeds and fails.
This is a creative research idea that deserves more exploration. I would love to see some quantification of the results here!
This is a very interesting direction. There seems to be some existing work around meta mech.interp so some more lit.review would help understand how this project fits into the current landscape. The first results seem promising
Cite this project
@misc{jacob2024let,
title = {{Let LLM Agents Perform LLM Surgery}},
author = {Sharat Jacob Jacob},
year = {2024},
month = nov,
note = {Submitted to Reprogramming AI Models Hackathon, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/let-llm-agents-perform-llm-surgery}},
url = {https://apartresearch.com/sprints/projects/let-llm-agents-perform-llm-surgery}
}More from Reprogramming AI Models Hackathon
- 1st place by peer reviewView project: AutoSteer: Weight-Preserving Reinforcement Learning for Interpretable Model Control
AutoSteer: Weight-Preserving Reinforcement Learning for Interpretable Model Control
AI Safety Initiative Groningen
Traditional fine-tuning methods for language models, while effective, often disrupt internal model features that could provide valuable insights into model behavior. We present a novel approach combining Reinforcement …
- View project: Classification on Latent Feature Activation for Detecting Adversarial Prompt Vulnerabilities
Classification on Latent Feature Activation for Detecting Adversarial Prompt Vulnerabilities
Feature Disruption Lab
We present a method leveraging Sparse Autoencoder (SAE)-derived feature activations to identify and mitigate adversarial prompt hijacking in large language models (LLMs). By training a logistic regression classifier on …
- View project: Utilitarian Decision-Making in Models - Evaluation and Steering
Utilitarian Decision-Making in Models - Evaluation and Steering
Byte To Brain
We design an eval based on the Oxford Utilitarianism Scale (OUS) that measures the model’s deontological vs utilitarian preference in a nine question, two factor model. Using this scale, we measure how feature steering …