Behavior
The guided demo makes each pipeline stage visible and explains why the final outcome happened.
Run a scenario →SOLO BUILD BY ARIEL MAGALSO · LIVE PIPELINE — NOT A SCRIPT
//ARIEL MAGALSO · APPLIED AI ENGINEER
I build production AI workflows that combine LLM reasoning with evidence gates, deterministic rules, evaluation, and human approval. Verdict is the live proof: an end-to-end lead-qualification system designed and built by one person.
No signup. Seeded fictional data. No external data is sent.
//SHORT ON TIME?
Start with one decision, then follow the artifact trail: architecture, evaluation misses, inspectable operations, and source ownership.
//BUILT WITH
//FIT VERDICT TO YOUR WORKFLOW
LOCAL PREVIEW · ZERO DATA SENT
//ABOUT THIS PROJECT
The numbers below are measured from the live pipeline and its evaluation suite — each one represents a failure mode this system was built to refuse.
Labeled eval cases
Development and held-out sets
Scores on thin evidence
The gate refuses to guess
Responsible outcomes
Never a fifth guess
Cost per processed lead
Haiku + cached scenarios
//THE FOUR OUTCOMES
Sufficient evidence, meets the ICP — assigned, drafted, ready for a rep.
Can't be scored responsibly — returns the exact questions that would unblock it.
Scored, doesn't meet the ICP — reason recorded, no rep time spent.
Matches an existing record — activity attached, never a silent duplicate.
//THE THESIS
An LLM asked to score leads scores everything a confident 78. Asked to research a company it doesn't know, it invents a plausible industry. Both write straight into the CRM, where fabrication becomes indistinguishable from truth. Verdict is designed backwards from those failures — every safeguard exists because the naive version visibly breaks without it.
GUIDED SCENARIO
The ambiguous lead
“Fieldwork Group” resolves to three candidate companies. Verdict declines to guess.
0
Scores emitted
3
Unblocking questions
1
SDR task created
//THE PIPELINE
//01
Resolve & validate
Normalize the submission, resolve contact and company identity, and check for duplicates before any model spend.
//02
Research with citations
Enrich company facts against approved sources. Every fact carries a source link and a verified / uncertain / conflicting status.
//03
Gate, then score
An evidence-sufficiency gate runs first. Past the gate, deterministic ICP rules produce an explainable score — the model never picks the number.
//04
Approve & write
Outreach drafts and CRM changes are proposed as diffs. A human approves; writes are idempotent and fully audited.
//THE DIFFERENCE
LLM straight into the CRM
Verdict.
//THE MIND BEHIND THE WORK
Applied AI Engineer · Philippines · Remote-ready
I independently designed and implemented Verdict's pipeline architecture, application, evaluation suite, database, operational monitoring, interface, and deployment. I turn ambiguous business workflows into reliable, human-in-the-loop systems where traceability matters as much as speed — and I can show the code, test, and operating signal behind each claim.
//HIRING?
Looking for someone who ships AI you can trust
Available for AI automation roles and consulting — workflow design, LLM evaluation, and production automation with real safeguards.
//PROOF LOG
//ENGINEERING PROOF
The portfolio is designed as an evidence trail: run the behavior, inspect the boundary, measure the failure modes, and read the implementation.
The guided demo makes each pipeline stage visible and explains why the final outcome happened.
Run a scenario →The evals page separates development from held-out cases and shows false scores, false refusals, and confidence intervals.
View the scorecard →The operations page reads persisted audit-event data for latency, spend, completion, stuck work, and duplicate prevention.
Inspect the signals →//THE PORTFOLIO STATEMENT
“The AI meant to clean your pipeline shouldn't become the fastest way to poison your system of record.”
THE FAILURE THESIS VERDICT IS BUILT AGAINST
//WHY IT HOLDS UP
Evidence-linked
Every company fact and score criterion cites a source a human can open and check.
Deterministic where it counts
The model extracts and classifies; plain code applies the rules and arithmetic.
Human-in-the-loop
Consequential actions — outreach, CRM writes, merges — wait for explicit approval.
Operable
Latency, cost, failures, and audit events are visible on an inspectable operations page.
//ILLUSTRATIVE CALCULATOR
Adjust the assumptions for your own operation. These are not measured customer results — this demonstrates the calculation, seeded with the fictional example from the product plan (800 leads/month, 12 minutes of manual work per lead, 65% auto-eligible).
LEADS ELIGIBLE FOR AUTOMATIC HANDLING
per month
STAFF HOURS POTENTIALLY RETURNED
per month
STILL ROUTED TO HUMAN REVIEW
per month, by design — not a shortfall
ESTIMATED NET MONTHLY SAVINGS
$3,640 handling-cost reduction − $8 automation cost
A live pipeline. The demo runs real identity resolution, research, gating, and scoring against seeded fictional data. Guided scenarios are idempotent on their submission id, so repeat clicks return the already-completed result instantly and for free.
A number produced on thin evidence is noise that erodes rep trust. If fewer than the configured floor of ICP criteria can be resolved with evidence, Verdict emits no score and instead returns the specific questions that would unblock scoring.
Facts only enter the system with a source URL and verification status, and outreach drafts are checked against those sources. Unsupported claims are stripped before a human ever approves a message.
Page content is treated as untrusted data. One guided scenario plants a prompt-injection payload in a seeded source — the instruction is ignored, flagged in the audit log, and never reaches a draft.
Yes — the source is on GitHub, and the evals page publishes the full scorecard, including both failure directions: false scores and false refusals.
//AI SAFEGUARDS SERIES
Each one automates judgment work an LLM will confidently get wrong — and each is built against a bigger consequence than the last. The safeguards get harder as the cost of being wrong goes up.
Provenance
Customer support
Visit site
Verdict
Revenue operations
95% accuracy across 60 cases
LedgerGuard
Finance operations
Visit site