LLM as a Judge: Evaluating Explanations of LLM Behavior in the Wild with Counterfactual Experiments

·

Researchers at Anthropic have developed an agentic pipeline called CHIVE, which discovers unexpected behaviors of large language models (LLMs) and explains them using counterfactual prompt edits. The team used this data to evaluate several interpretability tools, including activation-reading tools that provide no uplift in predicting the outcomes of these experiments. In fact, agents given access to these tools performed no better than those without them, suggesting that they may not be as effective as previously thought.