Skip to content
Sections
All notes

All notes · Monitoring

False Positives and What They Cost

Detection in this area generates a lot of noise. What that costs, who pays it, and how to keep the rate survivable.

Monitoring · Analysis

Pattern-based detection on prompt content produces alerts at a rate that surprises people. The handling cost, and the cost to the people flagged, both get underestimated.

The recommendations in “False Positives and What They Cost” become easier to sustain when implementation work has visible owners, dates and review time. Teams evaluating the detailed reference can use it to coordinate the operational side of AI adoption and identify where governance tasks are being missed, without treating activity data as evidence of misconduct or as a substitute for asking people why they chose a tool.

For an independent benchmark, compare the local approach with ICO guidance on AI and data protection; the useful test is whether ownership, access and recovery remain proportionate and explainable when the usual expert is absent.

Why the rate is high

AI prompts contain pasted material, which contains everything that was in the document.

Pattern matching on identifiers, reference numbers and formatted data fires on ordinary work.

And the context that would resolve it — this is a test record, this is already public — is not available to the matcher.

The analyst cost

Each alert needs a human look, and looking means reading what somebody wrote.

At volume this is both expensive and corrosive: the team reads a great deal of ordinary work to find the rare real case.

High-volume alerting also produces habituation, which is the documented failure mode in every detection discipline and ends with real alerts dismissed.

The cost to the person

Being flagged is unpleasant even when resolved.

If they are contacted, they learn that their prompts are read, and they will adjust — to personal devices, where you see nothing.

A programme generating frequent false positives teaches the organisation to avoid the monitored channel, which is the opposite of what it was built for.

Keeping the rate down

Narrow the patterns: specific identifier formats you actually hold, not generic ones.

Exclude test and sample data ranges.

Tune on real traffic for a period before acting on anything.

And set a threshold: single-field matches warn, multiple-field matches escalate.

The warn-rather-than-block question

Client-side warning — "this looks like client data, are you sure" — resolves most cases without anybody being flagged.

It also teaches, which blocking does not.

Blocking is appropriate for narrow categories where the answer is always no, and warning is better everywhere else.

Measuring it

Count alerts, and count how many were real.

If the precision is poor, the detection is not working regardless of how many true cases it also caught.

Publish that figure internally, because it is the thing that justifies tuning and it is never measured.

When to stop

If the rate cannot be brought to something the team can handle, the control is not viable as configured.

Narrow the scope rather than accepting a backlog nobody reads.

An alert queue with a backlog is worse than no alerting, because it provides false assurance.

What to check

What is your alert volume, and what proportion are real?

Does somebody read prompt content to triage them?

Do you warn at the client or block at the boundary?

And is there a backlog nobody is working through?

The point

A programme generating frequent false positives teaches the organisation to avoid the monitored channel, which is the opposite of its purpose..

Underlying all of this

Everything in this collection reduces to four habits: find out what people are doing and why before deciding anything, provide something good enough that the approved route is the easy one, write rules about information rather than about tools, and monitor the destination rather than the content. None requires a product, and a programme doing all four controls more than one built on prohibition.

The recurring pattern

The recurring pattern across every section here is the same: the response that feels like control reduces it. A ban removes visibility rather than use. Content inspection drives activity to personal devices. A discovery exercise with consequences produces quiet answers. In each case the organisation ends up knowing less about a risk it believes it has handled.