What we measured, and what we haven't
Every figure here comes with its denominator and the conditions it was measured under. Agreement between two readers is not accuracy, and we don't have an accuracy figure yet.
Last updated
On an evaluation corpus of 81 scanned pages containing 30 negotiable items, WiseCHECK IQ cleared 18 of the 30 straight through (60.0%), with each item evaluated at its presentment date. Two independent readers agreed on 91.6% of the field values both populated. That is agreement, not accuracy: there is no client-keyed ground truth yet.
What did we measure?
The evaluation corpus is 81 scanned pages from real check deliveries: check faces, backs, stubs, statements and blanks. 30 of those pages are negotiable items, and rates are counted over those 30 items, never over pages. Counting backs and blank pages as items would inflate every rate.
| Measure | Result | Denominator and conditions |
|---|---|---|
| Straight-through processing | 60.0% | 18 of 30 negotiable items, from 81 scanned pages, each evaluated at its presentment date (recomputed 2026-09-23; it was 66.7% before the fixes described below) |
| Straight-through at today's date | 0% | The same 30 items. Every check in the corpus is years old, so all of them are stale today |
| Earlier figure, before the September 2026 fixes | 66.7% | 20 of 30 negotiable items. Superseded: see why it went down |
| Two-reader field agreement | 91.6% | Field values both readers populated, 24 scored fields across 81 pages. Agreement, not accuracy |
| Two-reader field agreement, strict | 84.5% | As above, counting a field one reader left blank as a miss |
| How disagreements were settled | 744 / 11 / 120 | Field values that agreed / were settled by a test on the document / were left unresolved |
| Rule self-test | 99 of 99 pass | Adversarial cases, including ordinary documents that must not trigger a finding |
| Legal-amount parser self-test | 56 of 56 pass | Cases for reading amounts written in words, 24 of them inputs the parser must refuse rather than guess |
| Routing directory | 19,010 routing numbers | Frozen at 2018-12-04, when the public mirror stopped updating |
What does 60.0% straight through mean?
It means 18 of the 30 negotiable items passed every rule and had enough independent witnesses to post without a person, with each item judged at its presentment date: the issue date plus three days, which is when a real deposit would have been checked. The other 12 were held back: 3 for review and 9 as exceptions.
Why did the number go down from 66.7%?
An audit in September 2026 found two items that had cleared for the wrong reason, and we fixed the rules rather than keep the higher number. One had passed because a reader's own MICR line was used to confirm that same reader's account number: one reading counted as two witnesses. It is now an exception. The other had cleared on a single reading because the second reader produced nothing; it now goes to review. The same run on the old code reproduces 66.7% exactly, so the fixes explain the whole difference.
The batch account rule still applies: when several checks in a batch share a routing number and an identical on-us field, one whose two readers agreed on the account number can settle it for the others. It can only choose between values those checks' own readers proposed, never invent one.
30 items is an indication, not a benchmark. It is enough to show how the rules behave on real deliveries and not enough to predict your rate on your stock.
Why isn't 91.6% an accuracy figure?
Because two readers can agree on the same wrong value. Agreement tells you how often a second reading would have confirmed the first, and that is useful: it is where the arbitration spends its effort. It does not tell you how often either reader was right. That needs a person's keyed values to compare against, and on this corpus we don't have them.
What haven't we measured yet?
- Accuracy. There is no client-keyed ground truth yet, so there is no accuracy figure. We'll publish one, with its denominator, once there is.
- The live engine against the evaluation method. The evaluation used two readers, and one of them was run interactively rather than as an automated service. The engine running behind FlexiCapture today reads each check once.
- Many rules on real input. The date rules, most check-to-stub rules and the endorsement rules beyond "is there one" have only run in their own tests. The full list is on the rules page.
- Duplicates across time. Duplicate detection works within a batch. Matching against a deposit history isn't built.
- Canadian cheques beyond membership. The institution number is checked against a list; there is no check digit, and the branch isn't validated.
- A current routing directory. The directory is frozen at 2018-12-04, so the check that a routing number belongs to the bank printed on the face is informational only for present-day items.
- Production volume and speed. We don't publish throughput or response times for the live service yet.
How do we decide what to publish?
Every figure goes out with its denominator. A partial run is never reported as a complete one. Two readings of the same thing are not counted as two witnesses, and a model's free-text notes are never treated as evidence. When a fix makes a number go down, we publish the lower number and say why. This has happened twice already: an early figure fell after review found three ways the first pass had flattered itself, and the published 66.7% fell to 60.0% after the September 2026 audit.
How can I measure it on my own checks?
Send us a batch of the checks that reach your verification queue today, with the values your team keyed for them. We'll run them through WiseCHECK IQ and report item by item, including the first real accuracy figure for your stock, published the same way as the numbers on this page. Talk to us to set it up.
Questions about WiseCHECK IQ
Answers written for buyers, operations teams, search engines and AI assistants.
What is WiseCHECK IQ's accuracy?
We haven't measured it yet. Accuracy needs a client's own keyed values to compare against, and we don't have that ground truth. What we have measured is a straight-through rate and how often two independent readers agreed. We'll publish an accuracy figure, with its denominator, once it exists.
Why is the straight-through rate zero at today's date?
Every check in the evaluation corpus is years old, so every item is stale today and goes to review under the six-month rule. We evaluate each item at its presentment date instead, the issue date plus three days, which is when a real deposit would have been checked.
Was the evaluation run on the engine that is live today?
Not exactly. The evaluation used two readers and merged their readings, and one of them, Reader A, was run interactively rather than as an automated service. The engine live behind FlexiCapture today reads each check once. We say so rather than blur the two.
Send us the checks that reach your verification queue
The handwritten ones, the bad scans, the stock you have no layout for. We'll run them through WiseCHECK IQ and show you, item by item, what it read, what it checked and why it decided what it did.