How Investigators Handle Multilingual Evidence Without Losing the Meaning

Translation makes unreadable evidence readable, and quietly discards the part the case turns on.

Editorial image for How Investigators Handle Multilingual Evidence Without Losing the Meaning

A serious organised crime file in Europe is rarely monolingual. A single case can contain messaging threads in Arabic and Turkish, intercepted calls in a regional dialect that the national language model was never trained on, invoices in Mandarin, a device interface language in Russian, and an analyst team that reads none of it. The evidence exists. The ability to work it does not, until something translates it.

Machine translation has become good enough that this feels like a solved problem. Drop the thread in, get readable output, carry on with the case. The output is usually accurate enough to follow the conversation, which is exactly why the failure mode is easy to miss: the things translation loses are the things investigators build cases from, and they disappear without leaving a gap where the analyst can see one.

What translation removes that the case needed

Start with names. Transliteration is not deterministic, and the same person can arrive in a case file as four different spellings depending on which system handled which source. A name rendered from Arabic script into Latin characters by a phone extraction tool, by a bank's compliance system, and by a translation engine will frequently produce three variants. To a human they are obviously the same person. To any process that matches on strings, they are three entities, and the network graph splits into three unconnected clusters that each look too small to matter.

Then register and idiom. Coded language in a criminal conversation is not usually sophisticated, but it is local. A word for a quantity, a nickname for a border crossing, a euphemism for a payment method. Machine translation resolves these to their literal sense, which is generally the wrong one, and the output reads as an innocuous exchange about something domestic. The analyst sees a clean translation with nothing in it and moves on. The signal was in the choice of word, not the meaning of the word.

Then structure. Threads translated message by message lose the reply chains, the quoted fragments, and the timing of who was responding to whom. In a case where the argument is about coordination, the ordering carries as much weight as the content, and a flat translated transcript makes coordination hard to demonstrate even when the underlying data captured it precisely.

Why dialect is a harder problem than language

Most translation and transcription systems are evaluated on the standard written form of a language, which is not what appears in intercepted audio or in a messaging app. Spoken Arabic across North Africa and the Gulf varies enough that a model tuned to Modern Standard Arabic will mishandle a substantial share of a real call. The same applies to regional Hindi varieties, to spoken Chinese outside Mandarin, and to the mixed-language registers common in diaspora communities where two languages alternate inside a single sentence.

Transcription compounds it. A speech model that produces a fluent, grammatical transcript of an audio file it did not really understand is worse than one that produces an obviously broken transcript, because fluency reads as confidence. An analyst reviewing a clean transcript has no reason to request a human linguist. An analyst reviewing a visibly ragged one does. Systems that surface their own uncertainty are doing the investigator a service that systems optimised for polish are not.

This is why teams working high-volume multilingual intercept material almost always end up with a triage model rather than a translate-everything model. Machine output decides what a human linguist looks at. The linguist's work goes in the file. The machine output does not pretend to be evidence.

Keeping the original as the primary record

The discipline that separates teams who handle this well is simple to state and frequently broken in practice: the translation is a working aid, and the source language text or audio is the evidence. Every finding that rests on a translated passage should point at the original, and the file should make clear which is which.

The reason is not purism. It is that translations get challenged, and they get challenged successfully. A defence expert who can show that a key passage supports a different reading in the original language has not just weakened one finding, they have raised a question about every translated item in the file. When the original sits alongside the translation and the methodology is documented, that challenge is answerable. When the file contains only English output from a system nobody can name, it is not.

Practically this means the case system needs to hold both, linked, with the language of origin recorded and the means of translation recorded alongside it. Machine translated with which engine and which version, or translated by a named human, and reviewed by whom. That metadata costs nothing to capture at ingestion and is close to impossible to reconstruct a year later.

Correlating across languages rather than after translation

There is a structural choice buried in how a platform handles this, and it determines how much the language problem costs the investigation. One approach translates everything into a working language first, then correlates the translated text. The other correlates in the original languages, using representations that are not tied to a particular script, and translates only for human reading.

The difference shows up on exactly the cases that matter. If correlation happens after translation, every transliteration inconsistency and every mistranslated nickname is baked in before any linking occurs, and the entity that appears in Arabic in one source and in Latin script in another may never be recognised as one entity. If correlation happens across languages, a phone number is a phone number, an account identifier is an account identifier, and a person referenced in two scripts can be resolved as a single node with both surface forms recorded.

The practical test during an evaluation is short. Load a case containing the same individual named in two scripts, plus a document in a language nobody on the team reads, and ask whether the platform connects them. Then ask it to show why. A system that can produce the link and point at the two original source items has solved the part of the problem that matters. A system that produces fluent English and a graph with a hole in it has solved the part that was never the difficulty.

Request a Pilot

SentraLink correlates evidence across languages and scripts, keeps every finding pointed at its original source item, and runs its language models locally inside your environment.

Request a Pilot