WilWell Technologies
CivicAttest

← CivicAttest

CivicAttest Detection Accuracy Report — 40-Test Program

WilWell Technologies — Independent ALPR Usage Audit Tool Engine build r3 (2026-08-27), frozen throughout · r4 preview validated in parallel

The program

CivicAttest analyzes an agency's own exported license plate reader search logs and raises statistical patterns for human review. Its detection engine was measured against forty synthetic Flock-format audit logs of 5,000 records each — roughly 200,000 records — authored by an independent model (Perplexity) that never saw the tool's thresholds. Each log contained seven planted misuse patterns and deliberately planted "trap" scenarios: legitimate multi-officer incidents engineered to look statistically suspicious.

The program ran in three phases, each disclosed rather than blended:

Phase 1 — Development and confirmation (20 logs). Early exploratory tests exposed and fixed four detection defects. Ten development logs were then run blind and scored; one rule improvement followed. Six confirmation logs ran on the frozen engine: 32 of 42 patterns caught at row level (76%), 89% excluding one pattern whose generator contradicted its own answer key in every dataset examined.

Phase 2 — Adversarial red team (10 logs). The data author was asked to design evasion variants: gradual ramps under statistical gates, microbursts under the burst floor, low-volume external concentration, cross-agency paired searching, and vocabulary-masked off-hours activity. Three of seven evasion classes were still caught 10/10 (repeat-vehicle, masked off-hours, thin documentation at subject level). Four evaded as designed, and each became a named, specified fix.

Phase 3 — Standard persona-consistent series (10 logs). Fresh overt-pattern logs with a hard persona-consistency requirement, run on the frozen engine with results locked to disk before any answer key was opened.

Phase 3 headline results (frozen engine, 10 logs)

Planted pattern r3 (frozen series engine) r4.0 (shipped)
1. Repeated vehicle queries 10/10 10/10
2. Late-period volume increase 10/10 10/10
3. Off-hours by daytime account 10/10 10/10
4. Undocumented service-call searches 10/10 10/10
5. Solo overnight burst 10/10 — perfect row score all ten times 10/10
6. Outside-agency concentration 7/10 10/10
7. Broad network scope 2/10 10/10
All patterns 59/70 (84%) 70/70 (100%)*

59 of 70 patterns caught at pattern level (84%). Every one of the eleven misses is attributable to one of two specified causes with a validated fix (below). By design the engine flags enough of a pattern to trigger human review rather than every constituent row; at the row level this corresponds to 1,101 of 1,692 planted rows, and the review-lead posture is explained under Honest limitations.

False accusations — the program-wide number

Across all forty logs, containing roughly 2,400 deliberately planted trap rows representing legitimate stolen-vehicle and missing-person responses:

The accusatory rules — burst, off-hours, prohibited-purpose, and repeated-vehicle detection — touched zero trap rows. Zero, in forty tests.

Twenty-one trap rows across four logs were listed inside medium-severity volume-question flags, in every case because the officer genuinely worked the planted incident and their search rate genuinely rose during it — the flag's own "innocent explanation" field names exactly that cause. On fully clean random data, the engine produces zero findings about any individual.

The r4 release, validated before shipping

Every miss in the program maps to one of six specified fixes. One is already implemented and was validated live during Phase 3:

Network-scope rule (Pattern 7) — now shipped as build r4.0. Built in a parallel branch after test 3 of Phase 3, then required to predict Pattern 7's account on each remaining log before the answer key was opened. Ten for ten, including seven live pre-registered predictions, with zero false fires on clean data or any prior log. Following the series close it was promoted into both engines with unit tests, and two-engine parity was re-verified across all ten Phase 3 logs plus clean-data controls. With r4.0, Phase 3 pattern detection is 67 of 70 (96%).

Outside-account concentration rule and full-group case-rate evaluation — shipped as build r4.1. The concentration rule's thresholds came from measured separation, not tuning: organic outside accounts in the test data queried hundreds of distinct plates with a 0% repeat share, while every planted concentration sat at 80–100% repeat share on 4–6 vehicles — a margin any reasonable threshold clears. Retrospectively, r4.1 catches all 70 of 70 Phase 3 patterns (100%*) with no new trap contact anywhere in the program's forty logs and two-engine parity re-verified across 28 datasets.

* The retrospective 100% is marked as such because r4.1 shipped after Phase 3 answer keys were opened. A fresh frozen-engine confirmation batch is the standing protocol before this figure is quoted without the asterisk.

Remaining roadmap (r5), from the adversarial phase: microburst aggregation across dates, gradual-trend detection, and cross-agency pair detection.

Design properties verified during the program

Honest limitations


Test data authored independently without access to detection thresholds. All results locked to disk before answer keys were opened. The development/confirmation split, every inter-phase fix, and every miss are disclosed above. Reproduction: identical files through either engine yield identical findings.