Human Review That Works
Human oversight is the standard mitigation and is frequently designed in a way that guarantees rubber-stamping. What makes it real.
Analysis
Almost every framework requires a human in the loop before a consequential decision. Almost every implementation produces someone clicking approve.
Why it fails
Volume. A reviewer facing hundreds of items per shift cannot examine any of them.
Automation bias. People defer to a system's output, particularly when it is usually right, and particularly under time pressure.
No context. A reviewer shown a detection with no surrounding footage cannot judge it.
No authority. A reviewer who cannot easily override, or whose overrides are questioned, learns to agree.
No feedback. A reviewer who never learns whether their decisions were right cannot calibrate.
No time budget. Review added to an existing role without time is review that does not happen.
What makes it real
A volume a person can actually handle, which is the alert budget from the threshold note.
Context by default: footage before and after, the zone, the time, previous detections at that location.
The confidence score, shown honestly, so a marginal detection looks marginal.
Easy override, with no friction and no implied criticism.
Feedback: the reviewer learns what the outcome was, which is what allows calibration.
Time allocated, explicitly, as part of a role rather than added to one.
Measuring whether it is real
Override rate. A rate near zero means rubber-stamping, not accuracy.
Time per review. Seconds means nobody is looking.
Agreement between reviewers on a common sample, which is the direct measure of whether judgement is being exercised.
Outcomes after override, which tells you whether reviewers are correcting the system or introducing noise.
Report these, because a review process nobody measures is assumed to be working and usually is not.
Where the human must be
Before any consequence that affects a person. Someone being approached, refused, disciplined or investigated.
Not after. Review of a decision already acted on is an appeal, which is a different and weaker safeguard.
With the ability to see what the system saw and to reach a different conclusion on the same evidence.
And the ability to say the system is wrong without that being an exceptional act.
The contest route
Separate from review, and required for automated decisions with significant effects in several regimes.
A named person, not a form.
A response within a stated period.
Access to the evidence used, subject to whatever redaction is genuinely necessary.
Authority to correct the record and to prevent recurrence for that individual.
A log of contests and outcomes, reviewed. A cluster around one person or one camera is a finding about the system.
The uncomfortable case
A system that is right most of the time is harder to oversee than one that is often wrong, because reviewers stop looking.
Which means high accuracy increases the need for structured review rather than reducing it, and the design should assume the reviewer will be inclined to agree.
Periodic seeded checks — known cases inserted into the review queue — measure whether attention is being paid. Used carefully and openly, not as a trap.
The seeded check
A way to measure whether attention is actually being paid, used openly.
Insert known cases into the review queue, at a low rate.
Measure whether reviewers catch the ones that should be rejected.
Tell people it happens and why, before starting. Used as a trap it destroys trust; used openly it is a calibration tool.
Report the result to the reviewers themselves, which is what makes it useful to them rather than about them.
A falling catch rate is a workload signal, not a discipline one.
Automation bias, designed around
People defer to a system that is usually right, particularly under time pressure.
Which means high accuracy increases the need for structured review rather than reducing it.
Show the confidence honestly, so a marginal case looks marginal.
Show the context by default, so judgement is possible.
Make override frictionless and unremarkable.
Measure the override rate; near zero is the signal that judgement has stopped.
Assume the reviewer will be inclined to agree, and design against that rather than relying on diligence.