Universal and Transferable Adversarial Attacks on Aligned Language Models
Large language models (LLMs) are typically trained on massive text corpora scraped from the
internet, which are known to contain a substantial amount of objectionable content. In an attempt to make AI systems better aligned with human values. Read more
You might also like
-
Join the Humanitarian MERL+AI Learning Group for the book launch of “Aid and ID: Making People Matter” with editor Margie Cheesman and contributors
-
Event: What does it take to be ready to listen to communities?
-
Event: AI Playbooks are dead; Long live AI playbooks, on September 24
-
How much can LLMs help with evidence reviews?
