⏱️ Reading time: 12 min

Between February and July 2026, the British Transport Police (BTP) scanned more than 500,000 faces at London’s busiest stations using live facial recognition. The result: a single alert against the watchlist, which turned out to be a false positive, with no arrest directly resulting from it.

📑 En este artículo
  1. TL;DR
  2. What Is Facial Recognition?
  3. Why Accuracy Matters More Than the Headline
  4. How It Works: From a Face to a Vector of Numbers
    1. The Confusion Matrix: False Positives, False Negatives, and Their Cousins
    2. The Similarity Threshold: The Dial That Decides Everything
  5. Practical Python Examples
    1. Example 1: Distance Between Two Embeddings
    2. Example 2: Sweeping Several Thresholds
    3. Example 3: face_recognition with Explicit Tolerance
  6. Getting Started
  7. Real-World Use Cases
  8. Common Mistakes and Best Practices
  9. Comparison with Other Biometric Technologies
  10. Going Deeper: ROC, EER, and the Small-Database Problem
  11. Frequently Asked Questions
    1. What is the false positive rate in a face recognition technology?
    2. How is the similarity threshold chosen in a facial verification algorithm?
    3. Why did the facial recognition trial in London have only one false positive in 500,000 scans?
    4. What’s the difference between FAR and FRR in facial matching software?
    5. Is 1:1 verification the same as 1:N identification in a facial biometric identification system?
    6. Can I install facial matching software on my own computer?
  12. References

That number doesn’t prove the technology fails. It shows that, with a small and well-filtered watchlist, almost none of those 500,000 faces were ever going to match. Understanding why requires looking at how the accuracy of a facial biometric identification system is measured and what happens when the similarity threshold is moved.

TL;DR

  • The similarity threshold decides how many false positives and false negatives a facial recognition system produces.
  • FAR measures how many impostors a system accepts; FRR measures how many legitimate people it rejects.
  • Lowering the threshold reduces false negatives but drives up false positives, with no single point that eliminates both.
  • face_recognition, the Python library, uses a tolerance of 0.6 over 128-dimensional embeddings.
  • With a small watchlist and a strict threshold, 500,000 scans can yield just one false alert, as happened in London.

What Is Facial Recognition?

Facial recognition is a biometric technology that identifies or verifies people based on unique facial features, captured by a camera and converted into a numeric vector. It compares that vector against a database or watchlist, and decides whether there’s a match based on a configurable similarity threshold.

That last part almost never makes it into the headlines. The system doesn’t give an absolute yes or no: it calculates how similar two faces are and compares that number against a limit that someone, at some point, configured.

Why Accuracy Matters More Than the Headline

A headline like “500,000 scans, 1 false positive” can sound like a resounding success or a complete failure depending on who reads it. Neither interpretation is correct without two additional pieces of information: how many people were on the watchlist and how strict the threshold used was.

Fraser Sampson, former UK biometrics and surveillance camera commissioner, summed up the problem in comments reported by The Guardian: the success of a deployment depends on variables that need to align, including location, time of day, the composition of the watchlist, and the likelihood that those people will actually pass through that spot. Sampson, now a non-executive director at Facewatch (a company that operates facial recognition in stores), also pointed out a key difference: in a supermarket, success means known thieves don’t get in; for the police, success means catching someone. By that standard, the London trial wasn’t especially productive.

How It Works: From a Face to a Vector of Numbers

A camera captures a frame, a detection model locates the face and aligns it, and a convolutional neural network converts it into a vector of floating-point numbers, an embedding. The Python library face_recognition, based on dlib, generates 128-dimensional embeddings from that network.

flowchart TD
A["Camera captures the face"] --> B["Facial detection and alignment"]
B --> C["Neural network generates the embedding"]
C --> D["Comparison against the watchlist"]
D --> E{"Distance less than or equal to threshold?"}
E -->|"Yes"| F["Alert to officer"]
E -->|"No"| G["No match"]

That embedding isn’t stored as a photo: it’s a list of 128 numbers representing geometric features of the face (distance between the eyes, jaw shape, among others). Comparing two faces amounts to measuring how far apart their two vectors are in that 128-dimensional space.

The Confusion Matrix: False Positives, False Negatives, and Their Cousins

Each comparison against the watchlist falls into one of four categories: true positive (it was the right person and it matched), true negative (it wasn’t and it didn’t match), false positive (it wasn’t the person but the system said it was), and false negative (it was the person but the system said it wasn’t).

In biometrics, the false positive rate is called FAR (False Acceptance Rate) and the false negative rate, FRR (False Rejection Rate). NIST’s FRVT program measures and publishes both rates for dozens of algorithms from different vendors, instead of reducing them to a single “accuracy” number.

The Similarity Threshold: The Dial That Decides Everything

The threshold is the maximum distance the system tolerates before declaring a match. In face_recognition, that value is called tolerance and defaults to 0.6 over the Euclidean distance between two 128-dimensional embeddings.

Lowering the threshold (requiring more similar vectors) reduces FAR but raises FRR: the system becomes stricter and rejects more real matches. Raising it does the opposite. There’s no threshold that minimizes both errors at once, so every deployment chooses which error it would rather make.

sequenceDiagram
participant C as "Camera"
participant S as "Recognition system"
participant O as "Police officer"
C->>S: sends the frame
S->>S: calculates the embedding and distance
S-->>O: alert if distance is less than or equal to threshold
O->>O: verifies the match before acting
Note over S,O: the alert alone does not authorize an arrest
💡 Tip: if you need a single number to compare two systems, look for their Equal Error Rate (EER): the point where FAR and FRR are equal.
ThresholdEffect on false positives (FAR)Effect on false negatives (FRR)When to use it
Strict (e.g. 0.4)LowHighSurveillance with large watchlists or high cost of error
Medium (0.6, face_recognition’s default)BalancedBalancedPrototypes and general use cases
Permissive (e.g. 0.75)HighLowPersonal unlock, where convenience matters more

Practical Python Examples

Example 1: Distance Between Two Embeddings

Small vectors are enough to see the math; in production those vectors have 128 dimensions.

import numpy as np

embedding_a = np.array([0.10, 0.20, 0.30, 0.40])
embedding_b = np.array([0.15, 0.25, 0.28, 0.42])
threshold = 0.60

distance = np.linalg.norm(embedding_a - embedding_b)
print(f"Distance: {distance:.4f}")
print("MATCH" if distance <= threshold else "NO MATCH")

Output:

Distance: 0.0762
MATCH

Example 2: Sweeping Several Thresholds

This list simulates the distances of a face against five watchlist entries. Raising the threshold adds matches, whether real or false.

distances = [0.42, 0.58, 0.61, 0.73, 0.91]

for threshold in [0.4, 0.5, 0.6, 0.7]:
    matches = sum(1 for d in distances if d <= threshold)
    print(f"Threshold {threshold}: {matches} match(es)")

Output:

Threshold 0.4: 0 match(es)
Threshold 0.5: 1 match(es)
Threshold 0.6: 2 match(es)
Threshold 0.7: 3 match(es)

Example 3: face_recognition with Explicit Tolerance

import face_recognition

known = face_recognition.load_image_file("authorized_employee.jpg")
test = face_recognition.load_image_file("camera_capture.jpg")

known_encoding = face_recognition.face_encodings(known)[0]
test_encoding = face_recognition.face_encodings(test)[0]

result = face_recognition.compare_faces(
    [known_encoding], test_encoding, tolerance=0.6
)
distance = face_recognition.face_distance([known_encoding], test_encoding)[0]

print(f"Distance: {distance:.4f}")
print(f"Embedding length: {len(known_encoding)}")
print(result)

The exact distance depends on the photos you use, but two things are fixed with this library: the embedding always has 128 dimensions and compare_faces always returns a list of booleans. A typical output:

Distance: 0.4123
Embedding length: 128
[True]

Getting Started

On Linux (Ubuntu/Debian), install the build dependencies first, then the library:

sudo apt update && sudo apt install -y cmake build-essential python3-pip
pip install face_recognition

On macOS: brew install cmake && pip install face_recognition. On Windows: install Visual Studio Build Tools with C++ support, then run pip install cmake face_recognition.

To confirm it installed correctly, run:

python3 -c "import face_recognition; print('face_recognition installed successfully')"

Expected output:

face_recognition installed successfully
The BTP trial cost £320,786 in equipment and staffing across 18 deployments. Foto de Maxim Tolchinskiy en Unsplash

Real-World Use Cases

Dividing the trial’s total cost (£320,786) across the 18 deployments, each operation cost on average about £17,821 in equipment and staffing, for nearly 100 hours of total police work, according to the freedom-of-information request obtained by Liberty Investigates.

In August, the BTP extended the trial for four more months and expanded it to London Underground stations. Since that extension, police have reported three confirmed positive alerts: people found to be breaching sexual harm prevention orders or other court conditions. Unlike the first phase, the system got it right this time, which suggests the watchlist’s composition changed along with the operational threshold.

Transport for London justified the extension by saying the system would “identify people on police watchlists at key stations chosen for maximum impact,” with a focus on violence against women and girls. Meanwhile, more than half of the police forces in England and Wales already deploy live facial recognition, and the Metropolitan Police will install fixed cameras in London’s West End; Mayor Sadiq Khan confirmed its use on the newly pedestrianized Oxford Street.

Common Mistakes and Best Practices

  • Confusing few false positives with an accurate system: the absolute number depends on the watchlist’s size and the scanned volume, not just the threshold.
  • Using a single threshold for everything: airport security doesn’t carry the same cost of error as unlocking a phone.
  • Not versioning the threshold in production: changing it without logging the previous value breaks any later audit.
  • Ignoring demographic bias: the same threshold can produce a different FAR depending on demographic group if the model was trained on poorly balanced data, something NIST’s FRVT reports document per vendor.
  • Treating every alert as an automatic arrest instead of a lead that an officer must verify before acting.

Comparison with Other Biometric Technologies

TechnologyRequires contactWorks at a distanceTypical use case
FingerprintYesNoDevice unlock, access control
IrisNo, but requires proximityNoBorders, high-security facilities
VoiceNoYes, via microphoneCall centers, assistants
Facial identificationNoYes, via cameraSurveillance, unlock, border control

Going Deeper: ROC, EER, and the Small-Database Problem

Plotting FAR against FRR for every possible threshold value produces an ROC curve (or its DET variant, common in biometrics). Each point on that curve is a different threshold; the EER is the point where both rates cross, and it’s often used as a single-number summary when comparing algorithms, though it never replaces the full curve.

There’s an effect that explains, better than anything else, why London had just one false positive: if the watchlist has few people and the rest of the public has no connection to it, almost every comparison is a true negative by construction, so even a modest FAR produces few false positives in absolute terms. That same FAR, applied to a database of millions of faces (a national registry, for example), would generate far more false alerts without the algorithm having changed at all.

⚠️ Watch out: a threshold calibrated for a watchlist of a few hundred people can trigger many more false positives if applied to a database of millions of faces, even though the algorithm’s FAR hasn’t changed.
flowchart TD
A["Similarity threshold"] --> B["Stricter threshold, lower number"]
A --> C["More permissive threshold, higher number"]
B --> D["Low FAR, fewer false positives"]
B --> E["High FRR, more false negatives"]
C --> F["High FAR, more false positives"]
C --> G["Low FRR, fewer false negatives"]
More than half of England and Wales’s police forces already use live facial recognition. Foto de National Cancer Institute en Unsplash

Your next step: run Example 2 with your own simulated distances and note at which threshold your use case would rather be wrong, toward FAR or toward FRR.

📬 Get new articles by email

We only email about big articles (1-2 a month).

Frequently Asked Questions

What is the false positive rate in a face recognition technology?

It’s the FAR (False Acceptance Rate): the percentage of comparisons where the system declares a match with someone who is not actually the person being sought.

How is the similarity threshold chosen in a facial verification algorithm?

It’s chosen based on the cost of each type of error: a strict threshold reduces false positives but rejects more real matches, while a permissive threshold does the opposite. There’s no universally correct value.

Why did the facial recognition trial in London have only one false positive in 500,000 scans?

Because the watchlist used was small and controlled, so the vast majority of the 500,000 faces were true negatives by design, not because the algorithm was perfect.

What’s the difference between FAR and FRR in facial matching software?

FAR measures false positives (accepting an impostor) and FRR measures false negatives (rejecting the right person). Lowering one generally raises the other.

Is 1:1 verification the same as 1:N identification in a facial biometric identification system?

No: 1:1 verification compares a face against a single claimed identity (like unlocking a phone), while 1:N identification compares it against an entire watchlist, as in the BTP’s case.

Can I install facial matching software on my own computer?

Yes, the Python library face_recognition runs on Linux, macOS, and Windows with regular desktop hardware, with no need for a dedicated GPU for small test cases.

References

  • The Guardian: coverage of the British Transport Police’s facial recognition trial.
  • NIST FRVT: program that measures FAR and FRR for facial recognition algorithms from different vendors.
  • Wikipedia: history and general workings of facial recognition systems.
  • GitHub: face_recognition: the Python library used in this article’s examples, based on dlib.
  • OpenCV: official tutorial on face detection with cascade classifiers.

📱 Enjoy this content? Follow @programacion on Telegram for daily tech content in Spanish: quick summaries, fresh content every day. @programacion

Featured image: Foto de Miguel A Amutio en Unsplash

Did it work for you? Got a different error? Say so below: questions get answered and help the next reader.

Leave a comment
Categories: Tech NewsTutorials

Andrés Morales

Developer and AI researcher. Writes about language models, frameworks, developer tooling, and open source releases. Covers ML papers, the tech startup ecosystem, and programming trends.

0 Comments

Leave a Reply

Avatar placeholder

Your email address will not be published. Required fields are marked *

You can include code inside <code>…</code> or, for several lines, <pre><code>…</code></pre>.

This site uses Akismet to reduce spam. Learn how your comment data is processed.