A civic judge opens a case file at three in the morning. The system has already suggested a classification for the offense, drafted the facts and flagged a match with an earlier record. The judge reviews, accepts almost everything, changes one word and rules.
Six months later a complaint arrives. Someone asks why the offense was classified that way. The institution has the signed ruling, but it cannot answer the essential questions: did the system propose the classification or did the judge decide it? Did the judge check it against the medical certificate or accept it as it came? What did the original suggestion say before the change?
When a government adopts AI, the conversation centers on the model: how accurate it is, what data it was trained on, whether it is biased. Those questions matter. But in most public uses, AI does not decide on its own. It suggests, and a person decides. The quality of that decision depends on something the model does not control: whether there is a record of what the person did with the suggestion.
A suggestion is not a decision. And human oversight that leaves no trail cannot be demonstrated.
In public safety, civic justice and citizen services, the most common uses of AI replace no one. They draft the narrative of an event, propose the classification of an offense or a report, detect duplicates, compare a fingerprint against existing records or check a case file for inconsistencies.
In all of those cases there is a person at the end: the officer who signs, the judge who rules, the operator who accepts the classification. That is why these systems are said to have "a human in the loop," and why that presence is assumed to be enough of a safeguard.
Research on algorithm-assisted public decisions shows that assumption is incomplete.
Saar Alon-Barkat (University of Haifa) and Madalina Busuioc (Vrije Universiteit Amsterdam) published three experiments in the Journal of Public Administration Research and Theory on how algorithmic recommendations are processed in public decisions. The scenario: choosing which of three teachers on probation would not have their contract renewed, with a numeric score and a qualitative evaluation for each. Half of the participants were told the score came from an algorithm; the other half, from a human expert.
They looked for two biases: automation bias, following the system even when other sources contradict it, and selective adherence, following the recommendation when it confirms a stereotype.
Across the three studies, with 605, 904 and 1,345 participants, they did not find that people followed the algorithm more than the expert. But in the second study they did find selective adherence: the recommendation was followed more often when it matched a stereotype about the evaluated person's group, whether it came from an algorithm or an expert.
The third study was run with Dutch civil servants shortly after the childcare benefits scandal. There, the tax authority used an algorithm that, among other criteria, took nationality as a variable to flag high-risk applications, and agency staff manually reviewed the flagged applications. There was a person in the loop. Even so, the outcome disproportionately harmed families of foreign descent. The authors describe it as the meeting point between the system's bias and the reviewer's.
The lesson is direct: the presence of a person does not guarantee that the decision was reviewed. A suggestion can be accepted out of haste, workload or because it confirms what someone already thought. If the institution does not record what was suggested, what was accepted and what was changed, it has no way to tell a real review from an automatic signature.
The most advanced frameworks no longer settle for having a person involved. They require human intervention to be effective and documented.
The Government of Canada's Directive on Automated Decision-Making, in force since 2019 and updated in 2025, applies to any system that assists or replaces the judgment of a decision-maker in an administrative decision. It requires documenting the decisions made or assisted by the system, monitoring its outcomes to detect unfair effects, and recording client feedback, unexpected impacts and the times a person overrode the system, so those records can be used to correct problems.
The European Union's Artificial Intelligence Act points the same way. Its Article 14, on human oversight of high-risk systems, requires that those overseeing a system be able to understand its limits, interpret its outputs and disregard, override or reverse what it produces. It names automation bias explicitly: overseers must remain aware of the tendency to over-rely on the system, especially when it recommends decisions. And for certain biometric identification systems it requires that at least two competent people separately verify the identification before acting, with exceptions the regulation itself defines.
Mexican municipalities are not subject to these rules, but their logic works as a design standard: it is not enough for someone to decide; it must be possible to show what they decided in response to which suggestion.
The system writes its proposal straight into the final field. Afterwards there is no way to know whether that value was chosen or inherited.
The AI drafts the narrative of facts and the officer signs it with two edits. Months later no one can separate what the system wrote from what the person corrected, and that difference can matter in a complaint or a proceeding.
A fingerprint search suggests the detained person already has records under another name. The operator confirms, but only the final result is saved: not which candidates were shown, who confirmed or when.
Someone merges two records the system flagged as the same person. If it was a mistake, undoing it requires rebuilding by hand what came from each one, and if no one knows who approved it, the criteria cannot be reviewed.
If 99% of suggested classifications are accepted, the model may be excellent or no one may be reviewing them. Without acceptance and rejection data, the institution cannot tell the two scenarios apart.
Recording the human decision does not mean burying the operator in confirmation screens. It means the system stores, by design, five elements every time a suggestion touches a case file.
The value, text or match the system proposed, exactly as proposed, kept separate from the final value.
The person, their role and the moment they accepted, modified or rejected the suggestion, tied to that specific suggestion.
Accept it, modify it or discard it. At sensitive points, with a brief reason that puts the criteria in writing.
Whether the person had the medical certificate, the evidence or the history in front of them when deciding. Reviewing without the relevant information in view is not reviewing.
Acceptance, modification and rejection rates by type of suggestion, area and period. This is the data that turns individual records into institutional oversight: it shows whether human review works or has become a formality.
With those five elements, the complaint has an answer. The system proposed classification A; the judge reviewed it with the medical certificate in view, changed it to B and wrote down why. Or: the system proposed A and the judge accepted it unchanged in eleven seconds. Both answers are useful. The absence of an answer is not.
This is the other side of what we argued in Institutional AI fails when operations are not traceable: if operations leave no trail, AI learns from a weak record. If the decision about what AI proposes leaves no trail, the institution cannot oversee either the AI or the people who use it.
Tribuna, Intello's platform for public safety and civic justice, includes INTELLO AI, an assistant that helps draft the facts from the event, suggests the classification of offenses, checks the case file for consistency before it is sent and supports the structure of the medical report. The design principle is explicit: the assistant suggests and the decision belongs to the person.
That assistant runs on a foundation built to leave a trail. Tribuna's audit module records reads, writes and risk events, requires an access reason to open sensitive case files and keeps the full timeline of every case file. Every biometric comparison by fingerprint or face is logged. The person network detects duplicates by CURP, date of birth, address and phone number, and records are merged only with high confidence. And the judge rules with offenses, medical certificate, belongings and evidence in the same case file.
Agora, the citizen services platform, applies the same logic to a different kind of suggestion: smart deduplication and automatic classification of reports that arrive through web, app, chat, phone and walk-in counters. Its traceability keeps the full history of every case: who did what, when and with what result.
In both, AI is not a separate box: it works inside the same case file that is audited, and that is what makes human oversight demonstrable.
Before buying a tool that suggests, classifies or compares, it is worth asking:
- Does the system store the original suggestion separately from the final value?
- Is there a record of who accepted, modified or rejected each suggestion, and when?
- Does the person deciding have the relevant evidence in view, or only the suggestion?
- Do record merges and biometric confirmations leave a record of who approved them?
- Can the institution see acceptance and rejection rates by type of suggestion, area and period?
- For a specific case file, can you reconstruct what the system proposed and what the person decided?
If the vendor cannot answer clearly, the problem is not the model. It is that the institution will not be able to show it oversaw what the model did.
Most of AI's value in public operations will come from helping people work better, not from replacing them. But an assistant only strengthens an institution if every suggestion ends in a decision that can be traced: who made it, with what information and what they did with what the system proposed.
The question for leadership teams is not whether their system has a human in the loop. It is whether they can show, case file by case file, what that human did.
If your institution is evaluating how to adopt AI without losing traceability, see how Tribuna builds an assistant that suggests inside an audited case file, how Agora keeps the full history of every citizen case, or request a demo.