Topics I have worked on and still enjoy learning about.
I read brain science for its own sake, mostly perception and learning, and keep the notes in public. One finding has never stopped bothering me. People give the reasons for their choices fluently, at length, and are often wrong, which has been the standard result since the 1970s and is still the most inconvenient one. If that is what introspection is worth in us, what is a model's account of itself worth? A system asked to explain itself is producing a plausible account rather than reporting a mechanism, and the two can agree for a long time before they come apart. Behind all of it sits the question I actually like. When we say a system knows something, rather than that it reliably does something, what are we claiming?
Capability is rarely what limits a system built for people. The limit is the channel between them, and the channel is words. A request arrives as one sentence carrying a dozen unstated constraints, and what goes unsaid gets filled in from the training prior rather than from the person who asked. Philosophy of language calls the distance between what is said and what is meant implicature, and most of the work in a real request lives in that distance. Change how a question is framed and the judgment moves with it. We are the narrow part. I would rather work there than on another increment of capability.
What would have had to be different for the decision to go the other way? Of all the ways to explain a decision, this is the one that made sense to me. There are many true things you can say about a model, and nothing in the model tells you which of them is the explanation. A counterfactual at least has a shape. It hands you something to check instead of something to take on trust. The idea is much older than the methods that carry its name. It runs through Mill's method of difference, Lewis on counterfactual dependence, and Woodward's account of explanation as what would happen under intervention, and machine learning only took it up in the last decade. It is one method among several, and what recommends it is an argument rather than a benchmark.
Producing an explanation is the easy part. Judging whether it was any good has no floor under it. People want one before they can say what would satisfy them, and to score one you would need to know what the right explanation was, which is what you did not have. So the field measures proxies. Does the saliency map land inside the box a radiologist drew? That tells you how closely the model attends to what a person attends to. It does not tell you whether the map reflects what the model did. The distinction has hardened lately into two kinds of faithfulness. An explanation can reproduce a model's outputs while having nothing in common with its internal steps, or it can track those steps. Interpretability aims at the second, explainability usually at the first, and a good many disagreements about explainable AI turn out to be about which one was meant. Whether what is explained is a model or a person, the reasons for what it did have to be recoverable for the account to mean anything, and where the machine case falls short it falls short on means. If the methods improve they should move toward human explanation rather than away from it, because what an explanation has to satisfy is a person. That is also why nothing gets regulated that cannot first be explained.