Explainability Is Architectural: Why It Can't Be Added After Deployment

Updated: 1 day ago

The usual approach to explainable AI in financial services follows a predictable sequence. A model is built for performance. It is deployed. Then, when compliance, customers or regulators ask how it works, an explainability tool is attached to it: feature importances, SHAP values, counterfactual examples.
The result is an institution that can explain its model in great detail and still can't answer the question it is actually asked: why was this decision made?
That isn't a tooling problem. It is an architectural one. Explainability is not a layer that can be added after deployment. It is a property of how the decision system is designed, and if it isn't designed in, no tool can retrofit it.
What post-hoc explanation explains
Post-hoc explainability techniques are useful, and widely misunderstood. They answer a specific question: how did each input contribute to this model's output?
A SHAP analysis can show that a customer's credit score was lowered mainly by a short credit history and a high debt-to-income ratio. That is real information about the model.
But it leaves the decision unexplained. It doesn't say:
which threshold the score was compared to, and who set it;
which policy version applied;
whether other evidence, outside the model, influenced the outcome;
whether required evidence was missing;
whether an analyst intervened;
what the customer could change to obtain a different result.
The model's output was one input to the decision. Explaining it explains one link in a chain.
There is a second limitation. Post-hoc explanations are approximations. They describe the model's behavior locally, around one case, using a simplified representation. Different techniques can produce different explanations for the same prediction. That is acceptable for model development. It is fragile as the basis for telling a customer why they were denied credit.
The three audiences of an explanation
Explainability isn't one requirement. It is three, because three different audiences ask different questions.
Audience | Their question | What they need |
The analyst or operator | Can I trust this decision, and should I intervene? | The evidence, the signals and their uncertainty, the policy applied, what is unusual about this case |
The affected customer | Why was this decided about me, and what can I do? | The outcome, the criteria in plain language, the main factors, what would change the result, how to request review |
The regulator or auditor | Did the institution follow its own rules, consistently and fairly? | The full decision record, the policy version, consistency across equivalent cases, override patterns, reproducibility |
A model explanation partially serves the first audience. It barely serves the second and third. An architecture designed for explainability serves all three from the same source: the decision record.
What makes a decision explainable by design
A decision is explainable by design when its structure contains the answer to "why?" before anyone asks. Five architectural choices make that possible.
1. Separate the model from the policy. The model produces a signal: a probability, a score, an estimate with its uncertainty. An explicit, versioned policy turns that signal, together with other evidence, into a decision. When the two are separated, the decision can be explained in terms of the policy ("the estimated risk exceeded the threshold for this product"), and the model can be explained separately, if needed, to the people who need that level of detail.
When model and policy are fused, when "the policy" is whatever the model outputs, every explanation has to go through the model's internals, which customers can't understand and institutions often can't disclose.
2. Make the policy readable. Rules written in a form that people accountable for them can read can also be explained to the people affected by them. Rules buried in code can't.
3. Type the outcomes. When the outcome is one of a defined set (allow, deny, condition, escalate), each type can carry a defined explanation structure. A conditioned decision explains the condition. An escalated decision explains what triggered the escalation.
4. Record the evidence, including what was missing. An explanation that lists the factors considered but omits the evidence that was required and absent is incomplete. Missing evidence is often the most important part of the explanation.
5. Make decisions reproducible. An explanation is only trustworthy if the decision can be reproduced from its record. If the system can't regenerate the same outcome from the same inputs, the explanation is a story told after the fact.
Where language models fit
Language models are excellent at turning structured information into clear, natural prose. That makes them useful in explainable systems, with one strict boundary.
A language model can take a decision record and write an explanation for a customer in plain language, or summarize a complex case for an analyst. In that role, it translates something that already exists: the policy, the evidence and the outcome.
What it should never do is produce the reasoning itself. A language model asked "why was this loan denied?" without a structured decision record will produce a plausible answer, and plausibility is not truth.
Decisions can be explained in natural language. They should never be made, or justified after the fact, by a language model.
Counterfactuals that mean something
One of the most valuable things an explanation can offer a customer is a counterfactual: what would need to be different for the decision to change?
Model-level counterfactuals ("if your income were R$1,200 higher, your score would cross 0.7") are often misleading, because the score is not the decision. A policy-level counterfactual is grounded in the actual decision rule: "with updated financial statements, this request would be reassessed for the full amount" or "a limit increase is available once the current balance is below 80% of the existing limit."
Policy-level counterfactuals are actionable, honest and possible only when the policy is explicit.
Explainability and trade secrets
Institutions and vendors often argue that full explainability would expose proprietary models. When explainability is architectural, that tension largely disappears.
The explanation customers and regulators need is about the decision: the criteria, the evidence, the procedure. That can be shared in full without revealing a single model weight.
The model's internals remain protected, and can be examined by a supervisor under appropriate conditions if a specific concern arises.
A test for any AI decision system
Pick any decision your system made three months ago and try to answer these questions using only what the system recorded:
What was the outcome, and which policy version produced it?
Which evidence was used, and was any required evidence missing?
What did each model contribute, with what uncertainty?
Who had authority over the decision, and did anyone intervene?
What would have changed the outcome?
Can you reproduce the decision exactly?
Can you explain it in plain language to the customer, and in full detail to an auditor, from the same record?
If the answers require reconstruction, interviews or guesswork, the system is not explainable. It has explainability tools attached.
Designed in, or not at all
Explainability can't be bolted on after deployment for the same reason security and auditability can't: it depends on decisions made at the foundation of the system. How evidence is captured. How policy is represented. How outcomes are typed. How records are kept.
Institutions that design for it get explanations as a byproduct of how they decide. Institutions that don't will spend increasing effort producing explanations that never quite answer the question.
Explainability is not a feature of the model. It is a property of the decision architecture.
Want to see what decisions that are explainable by design look like? Talk to our team →

Comments