Case study: verified lead scoring
A reference implementation showing how a lead scoring and routing automation moves from SI-1 to SI-3 by adding a verifier, an autonomy ladder, and an audit log with rollback.
The problem
The original workflow scored leads from budget, urgency, and message sentiment, then routed them to the owner, head of sales, or a sales agent. Along the way it returned a total of zero from a wrong formula, errored on budget fields, and had a module that failed with no explanation. Nothing tested it, so each failure surfaced only after it had touched real leads.
What the reference adds
- Verifier
- Fifteen test leads with expected scores worked out by hand, plus rules that must hold for any input, such as "a bigger budget never lowers the score." No change reaches production unless all of it passes.
- Autonomy ladder
- Logging, alerts, and routing run alone. Customer replies and rule changes need human approval. Deleting leads is suggestion-only.
- Audit log and rollback
- Every decision is written to a tamper-evident log. Rules are versioned, and one command restores an earlier version.
- Fail loudly
- Unreadable or missing input goes to manual review with a reason. It never becomes a silent zero.
What the verifier caught
In the reference demo, an agent proposed a rule change containing a typo that zeroed the top budget tier. A human approved it. The verifier rejected it anyway, and production stayed on the previous rules. Approval alone is not verification.
Run it yourself
The Python reference implementation and a free-tier Make.com build guide are in the open repository.
The reference implementation meets the SI-3 evidence: a verifier with a test set, an autonomy ladder, an audit log, and rollback.