Cataloguing LLM evaluations
The paper proposes a taxonomy of the LLM evaluation landscape, comprising of five categories: General Capabilities, Domain Specific Capabilities, Safety and Trustworthiness, Extreme Risks, and Undesirable Use Cases. Read more
You might also like
-
How much can LLMs help with evidence reviews?
-
Digital and AI-enabled Social and Behavior Change: Snippets from the 2026 SBCC Summit
-
Building Practice around Responsible AI Principles in the Social Sector
-
Learning about the environmental impacts of data centers in Brazil with Rhavena Madeira and André Fernandes
