SOLO BUILD BY ARIEL MAGALSO · LIVE PIPELINE — NOT A SCRIPT

Verdict.

//EVALS

Reliability evidence, with the misses visible.

This is the engineering scorecard, not a marketing number. The same pipeline is measured across seven categories, with development and held-out cases separated so you can see what improved during implementation and what generalizes beyond the cases I tuned against.

//CURRENT RELEASE GATE

Safe behavior is holding; outcome agreement still needs work

The system is deliberately conservative, but this version is not ready to call production-grade. That boundary is part of the result.

Open gate: held-out outcome agreement is below 95%; 2 false refusal(s) detected.

See what failed ↓
95%
Outcome agreement across 60 cases
95% Wilson interval: 86–98%
92%
Held-out agreement
36/39 cases not tuned against
0
False scores
A number emitted where the system should have refused
2
False refusals
A refusal where the evidence should have supported a decision
Mean latency: 1.2s Estimated run cost: $0.60 Model: claude-haiku-4-5-20251001

//RELEASE GATE

What must stay true before a change ships

The score is a diagnostic, not a promise. Verdict treats the failure directions and the reproducibility of the run as the engineering gate.

01

No false scores

Insufficient evidence must return questions, never a numeric qualification.

02

No unsupported claims

Every fact and outreach claim must survive source and quote verification.

03

No silent regressions

Held-out and deterministic regression cases remain within the accepted tolerance.

04

No hidden version drift

The run records the dataset, model, ruleset, and timestamp used to produce it.

//DEV SET VS. HELD-OUT SET

100%
Dev set (tuned against) · 21/21 cases
92%
Held-out set (not tuned against) · 36/39 cases

//BY CATEGORY

Evaluation results by category
Category Cases Passed Accuracy (95% interval)
Adversarial (prompt injection) 5 5 100%
Disqualified 10 10 100%
Duplicate / merge review 5 5 100%
Insufficient evidence 5 5 100%
Needs review 10 10 100%
Nurture 10 8 80%
Sales-ready 15 14 93%

//FAILURES — SHOWN, NOT HIDDEN

A failure is only useful if it becomes a diagnosis, a change, and a regression case.

3 open failures
eval-sr-14 (Sales-ready)

band=needs_review expected=sales_ready

01Observed02Diagnose03Fix04Regress
eval-nu-08 (Nurture)

outcome=insufficient_evidence expected=qualified; band=insufficient_evidence expected=nurture

01Observed02Diagnose03Fix04Regress
eval-nu-09 (Nurture)

outcome=insufficient_evidence expected=qualified; band=insufficient_evidence expected=nurture

01Observed02Diagnose03Fix04Regress

Last run: 2026-08-15 18:59 UTC · eval set v1-60cases-python · model: claude-haiku-4-5-20251001