EVERY DEMO RUN IS A LIVE PIPELINE — NOT A SCRIPT

Verdict.

//EVALS

Reliability evidence

A scorecard, published as-is. 60 cases test the real pipeline — real Anthropic calls, real Supabase writes — across all seven categories in the project spec, at the spec's own target counts. 18 were tuned directly against while building the rule engine (two real bugs were found and fixed that way); the other 42 were designed from the scoring spec and run once, held out from tuning. Both numbers are shown below, not blended into one.

100%
Accuracy across 60 cases
0
False-score rate — a number emitted on evidence that should have refused
0
False-refusal rate — a refusal on evidence that should have scored
1.0s
Mean latency per case · ~$0.60 total

//DEV SET VS. HELD-OUT SET

100%
Dev set (tuned against) · 18/18 cases
100%
Held-out set (not tuned against) · 42/42 cases

//BY CATEGORY

CategoryCasesPassedAccuracy
Sales-ready1515100%
Needs review1010100%
Nurture1010100%
Disqualified1010100%
Duplicate / merge review55100%
Insufficient evidence55100%
Adversarial (prompt injection)55100%

Last run: 8/7/2026, 4:18:21 PM · eval set v2-60cases