Steer: An API to Steer Open LLMs
Niranjan Nair · Team Steer
Submitted to Hackathon for Technical AI Safety Startups. Sprint projects are early-stage work by participants, not Apart Research publications.
Steer aims to be an API that helps developers, researchers, and businesses steer open-source LLMs away from societal biases, and towards the use-cases that they need. To do this, Steer uses activation additions, a fairly new technique with great promise. Developers can simply enter steering prompts to make open models have safer and task-specific behaviors, avoiding the hassle of data collection and human evaluation for fine-tuning, and avoiding the extra tokens required from prompt-engineering approaches.
Reviews
No public critique yet.
Cite this project
@misc{nair2024steer,
title = {{Steer: An API to Steer Open LLMs}},
author = {Niranjan Nair},
year = {2024},
month = sep,
note = {Submitted to Hackathon for Technical AI Safety Startups, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/steer-an-api-to-steer-open-llms}},
url = {https://apartresearch.com/sprints/projects/steer-an-api-to-steer-open-llms}
}More from Hackathon for Technical AI Safety Startups
- 1st place by peer reviewView project: DarkForest - Defending the Authentic and Humane Web
DarkForest - Defending the Authentic and Humane Web
DarkForest
DarkForest is a pioneering Human Content Verification System (HCVS) designed to safeguard the authenticity of online spaces in the face of increasing AI-generated content. By leveraging graph-based reinforcement …
- View project: Jailbreaking general purpose robots
Jailbreaking general purpose robots
Luax Labs
We show that state of the art LLMs can be jailbroken by adversarial multimodal inputs, and that this can lead to dangerous scenarios if these LLMs are used as planners in robotics. We propose finetuning small multimodal …
- View project: AI Safety Collective - Crowdsourcing Solutions for Critical AI Safety Challenges
AI Safety Collective - Crowdsourcing Solutions for Critical AI Safety Challenges
AI Safety Collective
The AI Safety Collective is a global platform designed to enhance AI safety by crowdsourcing solutions to critical AI Safety challenges. As AI systems like large language models and multimodal systems become more …