Investigate arithmetic features in Multi-lingual LLMs
Akash Kundu, Ashish Rai, Suhas K R · Team one_dos_tres
Submitted to Reprogramming AI Models Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.
We investigate the arithmetic related feature activations in Llama3.1 70b model across its 8 supported languages. We use arithmetic-activation strength to compare the 8 languages and unsurprisingly English has the highest strength and Hindi, Thai score the least.
Reviews
This team finds some features that activate on GSM8K. They made an interesting decision to compare across languages.
With more time, I'd love to see this team investigate why they were unable to improve the performance of the model via steering.
Good initial work! Definitely interesting to see that different languages have lower activations in such a universal topic like maths. Would be interested to see the difference in language dependent and in-dependent features on other languages and math benchmarks.
Great research question. I find the intersection of math problems with its relatively clear evaluation criteria and multi-linguality a cool test bed for evaluating the robustness of feature steering. I hope the authors will iterate on this since it seems a worthy avenue!
Cite this project
@misc{kundu2024investigate,
title = {{Investigate arithmetic features in Multi-lingual LLMs}},
author = {Akash Kundu and Ashish Rai and Suhas K R},
year = {2024},
month = nov,
note = {Submitted to Reprogramming AI Models Hackathon, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/investigate-arithmetic-features-in-multi-lingual-llms}},
url = {https://apartresearch.com/sprints/projects/investigate-arithmetic-features-in-multi-lingual-llms}
}More from Reprogramming AI Models Hackathon
- 1st place by peer reviewView project: AutoSteer: Weight-Preserving Reinforcement Learning for Interpretable Model Control
AutoSteer: Weight-Preserving Reinforcement Learning for Interpretable Model Control
AI Safety Initiative Groningen
Traditional fine-tuning methods for language models, while effective, often disrupt internal model features that could provide valuable insights into model behavior. We present a novel approach combining Reinforcement …
- View project: Classification on Latent Feature Activation for Detecting Adversarial Prompt Vulnerabilities
Classification on Latent Feature Activation for Detecting Adversarial Prompt Vulnerabilities
Feature Disruption Lab
We present a method leveraging Sparse Autoencoder (SAE)-derived feature activations to identify and mitigate adversarial prompt hijacking in large language models (LLMs). By training a logistic regression classifier on …
- View project: Utilitarian Decision-Making in Models - Evaluation and Steering
Utilitarian Decision-Making in Models - Evaluation and Steering
Byte To Brain
We design an eval based on the Oxford Utilitarianism Scale (OUS) that measures the model’s deontological vs utilitarian preference in a nine question, two factor model. Using this scale, we measure how feature steering …