The forward-deployed engineer's defining tension: ship fast enough to keep the client's belief, and safely enough to keep their auditors calm. This demo is the resolution. Five acceptance criteria, committed and hashed before any result was seen. A labelled sample from the client's own history. A confidence threshold you must tune against the live accuracy-versus-coverage trade-off, a sign-off that unlocks only when everything passes, and a week-six drift re-run that pulls the approval the moment the world stops matching the sample. Runs live in your browser, link below.
The pilot dazzles, the steering committee applauds, and the AI goes live on the strength of a screen recording. Nobody wrote down what "good enough" meant, so nobody can say whether the system still is. Then the client wins a new contract, the vendor mix shifts, invoices arrive from a category the model has never seen, and the error rate climbs with no alarm attached to it, because the acceptance was an event and nobody made it a standard. I watched delivery under MAS Technology Risk Management for years at two financial institutions: banks solved this decades ago for conventional systems with UAT, acceptance criteria and re-certification. AI deployments deserve the same discipline, tightened for the fact that a model's competence is statistical and moves.
The five criteria are fixed in the data file and hashed on the page: overall accuracy on auto-approved invoices, minimum coverage (because a human queue that swallows 40 percent of volume has no business case), zero restricted-account invoices auto-approved at any confidence (a deterministic routing rule, verified rather than hoped), per-class floors, and reproducibility at production settings. You move the threshold and watch the two curves fight: left of the window, accuracy fails; right of it, coverage fails. Inside it, sign-off unlocks and binds your name to the operating point. From then on the drift tab owns you: the same five criteria re-run against week six's invoices, at your signed threshold, and the approval survives only if the evidence does.
At the slider's starting point (0.70) the machine takes 88 percent of invoices and codes 92 percent of them correctly: C1 fails, and so does the logistics class. Push right to 0.94 and accuracy is perfect on the third of invoices the model still dares to touch: C2 fails, because a two-thirds human queue has no business case. The window where everything passes is narrow, roughly 0.82 to 0.85, and finding it is the point: the operating point is a business decision wearing a statistics costume, and this page makes whoever sets it own it by name. Sign at 0.84 (96.0 percent accuracy, 75.7 percent coverage) and open week six: the client's new logistics contract doubled that vendor class, a renovation project introduced invoices the model has never seen, and at your signed threshold accuracy falls to 88.8 percent with the worst class at 78.9. C1 and C4 fail, the approval is suspended, and nobody had to notice, because the harness is the noticing.
| Operating point | Accuracy / coverage | Verdict |
|---|---|---|
| t = 0.70 (vendor default) | 92.0% / 88.0% | C1 + C4 fail |
| t = 0.84 (the window) | 96.0% / 75.7% | all five pass · sign-off unlocks |
| t = 0.94 (fear) | 100% / 34.3% | C2 fails · no business case |
| t = 0.84 · week-six sample | 88.8% / 74.2% | drift caught · re-acceptance required |
Where this stands, honestly. The model's proposals and the labels are synthetic, generated with roughly calibrated confidences and one deliberately weak vendor class, because what this demo proves is the harness: pre-commitment, visible trade-offs, gated sign-off, drift re-verification. It does not prove any real model's accuracy, and it says so on its evidence tab. In a live engagement the sample is the client's own labelled history, the criteria are negotiated with the process owner and internal audit before scoring, the reproducibility check runs at production settings (a model that answers differently twice cannot be accepted once), and the drift re-run is scheduled, not remembered. The same discipline my AuditBox build applies to AI-assisted bookkeeping applies here to the deployment itself.
Speed, with a spine. The forward-deployed model wins because someone competent is on the client's floor shipping weekly; it survives because what ships can face an auditor. This gate is the artifact that lets both be true at once: the business gets its Monday go-live, the risk function gets pre-committed criteria and a named signature, and the board gets a system that re-proves itself every time the world moves. The proposition in one line: I ship at field speed and produce the evidence trail an assurance partner would demand, because I trained on that side of the table.
committed 2026-07-14 · 5 criteria · sha-256 bound on page tuned t = 0.84 chosen against visible trade-off (96.0% / 75.7%) verified C1–C5 pass · restricted routing verified as rule, not threshold signed Reviewing engineer, named · go-live authorised week 6 drift sample: C1 88.8% · worst class 78.9% → C1, C4 FAIL outcome approval suspended · re-acceptance required · caught by design
I spent years delivering under MAS TRM, where nothing touches production without acceptance evidence, and a year at a global consultancy learning what auditors do to claims. Forward-deployed AI needs exactly that marriage: the pace of an engineer embedded at the client, and the paranoia of someone who knows the sample is not the season. So I pre-commit the criteria, make the trade-off visible enough that a business owner can genuinely choose, gate the signature on the evidence, and leave a harness behind that keeps asking the question after I have left the building. Fast is the job. Provable is the craft.
Start at the vendor default and watch two criteria fail. Find the window, sign at your chosen point, then open week six and watch the same harness take the approval back. The evidence tab shows every error the accuracy number is made of.