No false scores
Insufficient evidence must return questions, never a numeric qualification.
SOLO BUILD BY ARIEL MAGALSO · LIVE PIPELINE — NOT A SCRIPT
//EVALS
This is the engineering scorecard, not a marketing number. The same pipeline is measured across seven categories, with development and held-out cases separated so you can see what improved during implementation and what generalizes beyond the cases I tuned against.
//RELEASE GATE
The score is a diagnostic, not a promise. Verdict treats the failure directions and the reproducibility of the run as the engineering gate.
Insufficient evidence must return questions, never a numeric qualification.
Every fact and outreach claim must survive source and quote verification.
Held-out and deterministic regression cases remain within the accepted tolerance.
The run records the dataset, model, ruleset, and timestamp used to produce it.
| Category | Cases | Passed | Accuracy (95% interval) |
|---|---|---|---|
| Adversarial (prompt injection) | 5 | 5 | 100% |
| Disqualified | 10 | 10 | 100% |
| Duplicate / merge review | 5 | 5 | 100% |
| Insufficient evidence | 5 | 5 | 100% |
| Needs review | 10 | 10 | 100% |
| Nurture | 10 | 8 | 80% |
| Sales-ready | 15 | 14 | 93% |
A failure is only useful if it becomes a diagnosis, a change, and a regression case.
band=needs_review expected=sales_ready
outcome=insufficient_evidence expected=qualified; band=insufficient_evidence expected=nurture
outcome=insufficient_evidence expected=qualified; band=insufficient_evidence expected=nurture
Last run: 2026-08-15 18:59 UTC · eval set v1-60cases-python · model: claude-haiku-4-5-20251001