We deliberately did not let a model assign severity in a compliance-risk mapping output. That was not a judgment about whether the model was capable. It was a judgment about what the output was going to become.
The seductive option was obvious. A model could read inconsistent language, recognize nuance, and generate a severity or ownership suggestion much faster than a hand-maintained rule set. It could also attach a confidence score. But confidence is not replayability. If a result enters a governed record, another reviewer should be able to start with the same inputs, follow the same logic, and understand why that result exists.
That led to a separation by consequence.
First: candidate generation. This is where probabilistic assistance is useful. Let a model surface likely mappings, summarize patterns, identify anomalies, and narrow a large review surface. The output is explicitly a draft. It guides attention; it does not settle the record.
Second: record production. This is where the logic needs to be explicit. State the inputs. State the rules. Preserve the source. Make the transformation inspectable. Re-running the same inputs should reproduce the same output. That gives a reviewer something concrete to challenge: not a black box, but a rule that can be examined and changed.
Third: consequential judgment. A human remains accountable for exceptions, risk acceptance, approval, and any decision where the cost of error cannot simply be reversed. Human review is not decoration added after the model. It is a designed control boundary.
There is an important counterargument. Deterministic rules are not automatically better. They can be brittle. They can miss nuance. They can reproduce the assumptions and bias of their author flawlessly. A reproducibly wrong answer is still wrong. Replayability makes the method inspectable; it does not make the method correct.
So the useful question is not, can AI do this? Ask three narrower questions: What happens if the answer is wrong? Must another person be able to reproduce how the answer was reached? And who remains accountable when the output is accepted?
If the output only guides attention, probabilistic assistance may be appropriate. If the output becomes evidence, classification, or part of the governed record, make the logic replayable. If the output accepts consequence, keep a named human accountable.
The boundary is not between humans and machines. It is between suggestion, record, and decision. Design those three layers deliberately, and AI can add scale without quietly absorbing authority it was never meant to hold.
