Interpreting the Methods that Interpret Language Models
The Same AI, Four Different Explanations: Research Reveals Hidden Patterns Behind AI Explanations
AI is increasingly capable of making important decisions, but explaining why an AI model arrives at a particular decision proves to be less straightforward. Research by computational linguist Jonathan Kamp shows that different explanatory methods can provide different explanations for exactly the same AI decision. Moreover, these differences are not random: they often follow fixed patterns.
Artificial intelligence (AI) is increasingly being deployed for important applications, such as text evaluation, medical support, and detecting harmful online content. Therefore, methods have been developed to attempt to explain the 'reasoning' of an AI.
Kamp shows that this explanation is not always reliable. The differences between explanatory methods depend on the chosen method, the type of AI model, and the task the model performs. Furthermore, AI models can sometimes reach the correct outcome based on incorrect clues in a text.
Do Not Trust Blindly
By mapping the systematic patterns and biases in AI explanations, it becomes clearer when an explanation is reliable or, conversely, less reliable. The main conclusion is therefore that we should not simply trust AI explanations, but evaluate them critically.
The results are important for developers, companies, and users of AI systems who want to make reliable decisions. AI models must not only perform well but also be able to reliably explain why they make certain choices.
This is becoming increasingly relevant now that large language models such as chatbots are being used more frequently in education, healthcare, and information provision. For example, with an AI assistant providing medical information, it is possible to check which sources and text fragments influence the decision. This allows users to better assess whether the information is reliable.
More information on the thesis