Advanced English · Reading and VocabularyLesson 18 of 50

Lesson 18

Reasonable Degree of Certainty

Subject
Forensic science and the limits of expert comparison evidence
Register
Investigative-critical — measured, forensic cadence, precise about language
Level
C1–C2
Extent
1,173 words

An expert takes the stand. They describe their training, their years of practice, the number of examinations they have conducted. They are asked whether, in their professional opinion, the mark recovered from the scene was made by the defendant. They say that it was, and they add a phrase that has been repeated in courtrooms for the better part of a century: to a reasonable degree of scientific certainty.

The phrase has no agreed definition. It is not a statistical statement, it corresponds to no threshold or confidence interval, and it originates in legal practice rather than in any laboratory. It sounds, to a jury, like a measurement. It is a form of words.

This is the appropriate place to begin, because the central problem with comparison forensics is not fraud, and it is only partly method. It is that a set of practices developed by investigators, for investigative purposes, acquired the vocabulary and the authority of science without ever acquiring its infrastructure of validation.

What a national review found

In 2009, a review commissioned from the United States National Academy of Sciences examined the forensic disciplines as a whole. Its central finding was stated without hedging: with the exception of nuclear DNA analysis, no forensic method had been rigorously shown to have the capacity to demonstrate consistently, and with a high degree of certainty, a connection between evidence and a specific individual or source.

The finding applied to hair comparison, tool marks, bite marks, handwriting, fibre analysis, and — most consequentially, because it is the most trusted — friction ridge comparison, which the public knows as fingerprinting.

A further review in 2016 assessed which feature-comparison methods had actually been subjected to properly designed studies establishing their accuracy. The list was short. For several long-established disciplines, appropriately conducted validation studies did not exist at all. Not studies with poor results: studies of the necessary kind had not been performed.

How this happened

Understanding the history makes the situation less surprising and more troubling.

Most of these techniques were developed inside police laboratories from the late nineteenth century onwards, by practitioners solving immediate cases. They worked, in the sense that they produced leads and the leads were often correct. Their validity was established by use rather than by experiment, and by the time anyone asked for measured error rates, the methods had been admitted in courts for decades, and their admissibility rested substantially on the fact of prior admission.

Medicine passed through the same phase and left it, under pressure, in the second half of the twentieth century — abandoning treatments that had been used confidently for generations once controlled trials showed they did not work. Forensic comparison had no equivalent reckoning, largely because its practitioners were not in competition with each other for publication, and because the institution evaluating their claims was a court rather than a journal.

The disciplines that failed

Two examples establish the range.

Bite mark comparison — matching an injury on skin to a suspect's dentition — was accepted in courts for decades and used in serious cases. Subsequent testing was devastating. Skin is elastic, deforms on contact, and changes over hours. Controlled studies found that examiners disagreed with each other at high rates, and frequently disagreed about the prior question of whether an injury was a bite at all. A substantial number of convictions in which such testimony was central have since been overturned, several after the convicted person had served decades.

Microscopic hair comparison was subjected to internal review by a major federal laboratory, examining cases in which examiners had testified. The review found that in the large majority of the cases assessed, the testimony had contained statements exceeding what the science supported — typically by describing a degree of association between a questioned hair and a specific individual that microscopy cannot establish.

Neither review concluded that the examiners had lied. They had testified to a confidence the discipline had never earned and which nobody had previously required them to justify.

The harder case: fingerprints

Friction ridge comparison is a genuinely powerful technique, and it should not be equated with bite marks. Ridge patterns are highly variable, they are persistent through life, and skilled examiners performing careful comparisons are demonstrably accurate. Large studies have produced false positive rates on the order of one in a thousand — low, meaningfully measured, and not zero.

Three problems remain.

The first is that the discipline long maintained that its error rate was zero and that identification was absolute. This position was untenable and was abandoned only under external pressure, and its residue persists in how conclusions are phrased to juries.

The second is subjectivity in a specific place. The method has no defined numerical threshold — no minimum number of corresponding features required for a conclusion. The examiner judges sufficiency. That judgement is informed by long experience and is not therefore arbitrary, but it cannot be audited, because there is no stated standard against which to audit it.

The third is context. Experiments in which examiners were shown prints they had previously analysed, accompanied by biasing information about the case, found that a proportion changed their conclusions. This is ordinary human cognition operating as it always does, and it is precisely why other empirical disciplines blind their analysts. Forensic examiners have historically worked inside investigative agencies, receiving case information as a matter of routine, because nobody had identified this as a hazard.

A widely reported misidentification following a major bombing investigation demonstrated all three at once. A print was attributed to a man with no connection to the events. The conclusion was confirmed by additional examiners and by an independent expert appointed for the defence. It was wrong. The subsequent review identified circular verification — later examiners knowing the first conclusion — and contextual pressure as contributing causes.

What reform requires

The technical remedies are neither exotic nor expensive, and they are largely borrowed from other fields.

Verification should be blind: the second examiner should not know what the first concluded. Analysts should receive only the information required for the comparison, released in sequence, with domain-irrelevant case detail withheld. Error rates should be measured in realistic conditions and disclosed in testimony. Conclusions should be stated in terms of the strength of evidence supporting one proposition over another, rather than as declarations of identity. And laboratories should be administratively independent of the investigating agency, because a unit whose funding and standing derive from the body it is evaluating evidence for is not structurally positioned to disappoint it.

Progress has been real but uneven, and the disciplines have generally moved fastest where an appellate court has forced them to.

The point is narrower than the popular framing of scandal. Nothing here shows that forensic evidence is worthless; some of it is extremely strong. It shows that the confidence expressed in courtrooms has for a long time been calibrated by professional custom rather than by measurement — and that a jury cannot discount an overstatement it has no way of recognising as one.

Key vocabulary

take the stand v. phr.
to enter the witness box to give evidence.
defendant n.
the person accused in a legal case.
threshold n.
a defined level at which a conclusion becomes permissible.
validation n.
formal testing establishing that a method performs as claimed.
hedging n.
the use of cautious, qualifying language.
friction ridge n. phr.
the raised skin pattern on fingers and palms.
admissibility n.
the legal quality of being permitted as evidence.
reckoning n.
a settling of accounts; a moment of critical assessment.
dentition n.
the arrangement of a person's teeth.
elastic adj.
able to deform and return to shape.
overturn v.
to reverse a legal decision.
questioned adj.
of an item, the one of unknown origin under examination.
untenable adj.
impossible to defend against objection.
residue n.
what remains after most of something has gone.
sufficiency n.
the judgement that there is enough of something.
audit v.
to check formally against a defined standard.
blind v.
to withhold information that could bias a judgement.
circular adj.
where a conclusion depends on what it is meant to confirm.
attribute to v. phr.
to identify as originating from.
appellate adj.
concerned with the review of lower-court decisions.
calibrate v.
to set against a reference standard.
discount v.
to reduce the weight given to something.

Phrases and collocations

a form of words
a phrase with the appearance of meaning but no defined content.
without hedging
stated plainly, with no qualifying language.
rest on
to depend for support upon.
exceed what the science supports
to claim more than the evidence permits.
as a matter of routine
habitually, without any decision being taken.
structurally positioned to
placed by institutional arrangement so as to be able to.