Units that have adopted correlation tooling in the last two years have mostly worked out how to get findings out of it. Far fewer have worked out what happens between a finding appearing on screen and that finding appearing in a document that goes to a prosecutor. In many teams the answer is that an analyst read it, thought it looked right, and copied it across.
That gap is where the risk concentrates. Not in the model, which is usually more conservative than the people using it, but in the absence of a defined step where a human takes ownership of the claim. Agencies that have been through a challenged case tend to build that step deliberately afterwards. It is considerably cheaper to build it first.
Three kinds of output, three different burdens
Treating all machine output the same way is the first mistake. There are three distinct categories and they need different handling.
Extraction is the system reading something and reporting what it says: an account number on a statement, a date on an invoice, a name in a chat message. Verification is direct. Open the source, look at it, confirm the characters match. It takes seconds and it is the one category that can be checked with certainty.
Correlation is the system asserting that two things in different sources refer to the same entity or event. This cannot be verified by looking at one item, because the claim is about a relationship. Verification means examining both sources and deciding whether the identifiers genuinely match or merely resemble each other. Most false findings live here, and most of them are near-misses rather than nonsense: two accounts at the same institution with adjacent numbers, two people who share a surname, two events an hour apart that the system treated as simultaneous.
Inference is the system characterising what the evidence means. This one carries the heaviest burden, because the output is a conclusion and conclusions in a case file belong to people. The analyst is not checking whether the system was right. The analyst is deciding, on the evidence, what they themselves conclude, and taking responsibility for it.
What review has to leave behind
Review that produces no record is indistinguishable from review that never happened, which is the position a lot of teams discover themselves in when a case is questioned two years later.
A finding that has been properly reviewed should carry four things. What the system asserted, in its original form rather than paraphrased into the report. Which source items it drew on, addressable so a third party can open them. Who reviewed it, and when. And what the reviewer concluded, including where they disagreed with the system or narrowed its claim.
That last element is the one most often dropped and the most valuable. A file showing that an analyst examined nine correlations, accepted six, rejected two as coincidental, and narrowed one to a weaker claim is a file that demonstrates method. It answers the question a defence expert is going to ask, which is not whether the software works but whether anybody was actually looking.
None of this needs to be onerous. It needs to be part of the tool rather than a parallel spreadsheet, because a review log maintained separately from the case system diverges from it within weeks.
The failure modes that survive review
Some errors pass review reliably, and knowing them is most of the defence against them.
Plausibility bias is the main one. A correlation that fits the theory of the case gets less scrutiny than one that contradicts it. The countermeasure is procedural rather than personal: review findings before deciding they support the theory, and give the ones that fit best the closest look rather than the quickest.
Volume fatigue is the second. Verification quality falls off measurably after a few dozen items in a sitting, and the last twenty in a batch of two hundred get a fraction of the attention the first twenty received. Teams that split review across sessions and across people catch materially more.
Derived-finding drift is the third and the most technical. Analyst A verifies a correlation. Analyst B builds a conclusion on top of it. Then the underlying evidence is withdrawn, or reclassified, or found to have been misattributed. If the system does not propagate that change to everything derived from it, the conclusion stays in the file with its foundation removed, and nobody notices until someone goes looking. Any platform used for serious casework should be able to answer the question: what in this file depends on this item.
Writing it down before you need it
The teams handling this best have a short written standard, usually a page or two, and it exists before the first challenged case rather than after. It says which categories of output require verification against source and which can be accepted with a spot check. It says who is authorised to sign off a finding for inclusion in a report. It says what happens to derived findings when source evidence is withdrawn. It says how machine assistance is described in the report, in language the unit has agreed rather than language each analyst invents.
That last point deserves emphasis. Reports vary wildly in how they describe tooling, from silence to language that overstates what the system did. Silence invites the accusation that the agency concealed its methods. Overstatement hands the defence a target. The workable middle is plain description: the analyst used the platform to identify candidate correlations across the listed evidence sources, examined each against the source material, and reached the stated conclusions.
That sentence is accurate, it is defensible, and it puts the finding where it belongs, with the investigator who made it.