Running a Pilot That Can Fail
A pilot designed to succeed is a procurement exercise. What to fix in advance so the result means something.
Procedure
Most pilots in this field conclude that the system works, because they were structured so that no other conclusion was available.
Decide the criteria first
Written down, before the pilot starts.
The minimum recall, with a number.
The maximum false alarms per camera per day.
The conditions it must work in, including the difficult ones.
Who decides, and on what evidence.
What result would mean no. If nothing would, the pilot is decoration and everyone's time is better spent elsewhere.
Structure it honestly
Your cameras, your site, your conditions.
Long enough to include the hard periods: the dark months, the busy season, the weather.
Ground truth collected independently, by someone recording what actually happened without seeing the system's output.
No vendor tuning during the measurement period, or you are measuring their attention rather than the product.
A separate tuning phase before it, with the measurement phase clearly bounded.
Ground truth is the expensive part
And it is the part that gets cut, which is why so many pilots prove nothing.
A person reviewing footage independently and recording events against the agreed definition.
Both directions: what the system reported, and what actually occurred including what it missed.
Sampled if full coverage is impractical, with the sampling method recorded.
Budget for it explicitly. Without ground truth you have a demonstration with a longer duration.
The people question during a pilot
A pilot involving people needs the same basis, notice and assessment as a deployment.
"It is only a trial" is not an exemption, and treating it as one is the most common compliance failure in this field.
Tell people it is running, what it does, and for how long.
Delete the pilot data at the end unless there is a stated reason to keep it.
Reporting the result
The measured numbers, at the stated threshold, against the criteria.
By condition, so a system that works in daylight and fails at dusk is described accurately.
By subgroup where applicable.
The alert volume that would result at full deployment, which is the arithmetic that most often changes the decision.
What was not tested, which is the honest section.
Concluding no
A pilot that fails its criteria has succeeded as a pilot.
Say so plainly, with the numbers.
Distinguish the causes: the product, the camera placement, or the application being unsuitable. These lead to different next steps and conflating them wastes the finding.
Camera placement is the most common cause and it is fixable, which makes it worth separating from a product judgement.
Record the decision and the evidence, because the same product will be proposed again in two years and the file is the answer.
Separating the causes of a failed pilot
A pilot that fails is useful only if the cause is identified.
The camera: wrong angle, insufficient pixels on target, bad light. Fixable, and the most common cause.
The product: genuinely underperforms on this task in these conditions.
The application: the event is not definable, too rare, or not visually distinctive.
Test the first before concluding the second, by moving one camera and re-running.
Record which it was, because the same product will be proposed again and the file should say whether it was the product or the mounting.
The pilot is not exempt
The most common compliance failure in this field.
A pilot involving people needs the same basis, notice and assessment as a deployment.
"It is only a trial" is not an exemption in any regime.
Tell people it is running, what it does, for how long.
Delete the pilot data at the end unless there is a stated reason to keep it.
Record the pilot in the register while it runs, and close the entry when it ends — which is also how pilots stop becoming permanent by inattention.