Pinned
Thanks for sharing, @_akhaliq!
We study how an adversary can *exploit* instruction tuning via data poisoning.
For example, one can inject training data that promote their products in the example responses, and we find that the model can pick up this behavior.
On the Exploitability of Instruction Tuning
paper page: huggingface.co/papers/2306.17…
Instruction tuning is an effective technique to align large language models (LLMs) with human intents. In this work, we investigate how an adversary can exploit instruction tuning by injecting



