WTWilly Tai
Case study · live demoForward-Deployed Engineering · Verification

Everyone wants it live on Monday. This gate is what makes Monday safe.

The forward-deployed engineer's defining tension: ship fast enough to keep the client's belief, and safely enough to keep their auditors calm. This demo is the resolution. Five acceptance criteria, committed and hashed before any result was seen. A labelled sample from the client's own history. A confidence threshold you must tune against the live accuracy-versus-coverage trade-off, a sign-off that unlocks only when everything passes, and a week-six drift re-run that pulls the approval the moment the world stops matching the sample. Runs live in your browser, link below.

01 / The pain

Deployments go live on demo energy, and quietly rot from week four.

The pilot dazzles, the steering committee applauds, and the AI goes live on the strength of a screen recording. Nobody wrote down what "good enough" meant, so nobody can say whether the system still is. Then the client wins a new contract, the vendor mix shifts, invoices arrive from a category the model has never seen, and the error rate climbs with no alarm attached to it, because the acceptance was an event and nobody made it a standard. I watched delivery under MAS Technology Risk Management for years at two financial institutions: banks solved this decades ago for conventional systems with UAT, acceptance criteria and re-certification. AI deployments deserve the same discipline, tightened for the fact that a model's competence is statistical and moves.

02 / Why the naive approach fails

"It scored 96% in testing" is not an acceptance. It is an anecdote with a number.

03 / How it works

Commit, tune, prove, sign. Then keep proving.

The five criteria are fixed in the data file and hashed on the page: overall accuracy on auto-approved invoices, minimum coverage (because a human queue that swallows 40 percent of volume has no business case), zero restricted-account invoices auto-approved at any confidence (a deterministic routing rule, verified rather than hoped), per-class floors, and reproducibility at production settings. You move the threshold and watch the two curves fight: left of the window, accuracy fails; right of it, coverage fails. Inside it, sign-off unlocks and binds your name to the operating point. From then on the drift tab owns you: the same five criteria re-run against week six's invoices, at your signed threshold, and the approval survives only if the evidence does.

criteria committedhashed before results threshold tunedtrade-off visible · yours to own 5 criteria passonly inside the window named sign-offgo-live authorised drift re-runsame criteria, forever fails → approval suspended, back to the gate acceptance is an event · drift is a season · the same arithmetic governs both
fig.1 · the gate is a loop, not a ceremony: what passed must keep passing
04 / One worked example

The window is 0.82 to 0.85, and week six breaks it on schedule.

At the slider's starting point (0.70) the machine takes 88 percent of invoices and codes 92 percent of them correctly: C1 fails, and so does the logistics class. Push right to 0.94 and accuracy is perfect on the third of invoices the model still dares to touch: C2 fails, because a two-thirds human queue has no business case. The window where everything passes is narrow, roughly 0.82 to 0.85, and finding it is the point: the operating point is a business decision wearing a statistics costume, and this page makes whoever sets it own it by name. Sign at 0.84 (96.0 percent accuracy, 75.7 percent coverage) and open week six: the client's new logistics contract doubled that vendor class, a renovation project introduced invoices the model has never seen, and at your signed threshold accuracy falls to 88.8 percent with the worst class at 78.9. C1 and C4 fail, the approval is suspended, and nobody had to notice, because the harness is the noticing.

Operating pointAccuracy / coverageVerdict
t = 0.70 (vendor default)92.0% / 88.0%C1 + C4 fail
t = 0.84 (the window)96.0% / 75.7%all five pass · sign-off unlocks
t = 0.94 (fear)100% / 34.3%C2 fails · no business case
t = 0.84 · week-six sample88.8% / 74.2%drift caught · re-acceptance required
C3 never appears in the trade-off, on purpose. Restricted-account invoices route to a person at any confidence, by deterministic rule rather than by threshold. Some risks are not statistical questions, and the harness's job is to verify the rule is wired, not to price it.
05 / Numbers & honest trade-offs

What it does, and what it does and doesn't prove.

5
pre-committed criteria
Hashed on the page; change a word, the hash changes.
300
labelled acceptance invoices
Plus 120 in the week-six drift sample.
0.82–0.85
the passing window
Found by you, owned by you, signed by you.
2
criteria broken by drift
At the signed point; approval suspended automatically.

Where this stands, honestly. The model's proposals and the labels are synthetic, generated with roughly calibrated confidences and one deliberately weak vendor class, because what this demo proves is the harness: pre-commitment, visible trade-offs, gated sign-off, drift re-verification. It does not prove any real model's accuracy, and it says so on its evidence tab. In a live engagement the sample is the client's own labelled history, the criteria are negotiated with the process owner and internal audit before scoring, the reproducibility check runs at production settings (a model that answers differently twice cannot be accepted once), and the drift re-run is scheduled, not remembered. The same discipline my AuditBox build applies to AI-assisted bookkeeping applies here to the deployment itself.

06 / The value proposition

What a client buys when they buy this.

Speed, with a spine. The forward-deployed model wins because someone competent is on the client's floor shipping weekly; it survives because what ships can face an auditor. This gate is the artifact that lets both be true at once: the business gets its Monday go-live, the risk function gets pre-committed criteria and a named signature, and the board gets a system that re-proves itself every time the world moves. The proposition in one line: I ship at field speed and produce the evidence trail an assurance partner would demand, because I trained on that side of the table.

Acceptance record · signed, then challengedcriteria hash bound before scoring
committed  2026-07-14 · 5 criteria · sha-256 bound on page
tuned      t = 0.84 chosen against visible trade-off (96.0% / 75.7%)
verified   C1–C5 pass · restricted routing verified as rule, not threshold
signed     Reviewing engineer, named · go-live authorised
week 6     drift sample: C1 88.8% · worst class 78.9% → C1, C4 FAIL
outcome    approval suspended · re-acceptance required · caught by design
The same shape as every instrument on this site: deterministic rules decide, evidence accumulates, a named human owns the consequential call, and the system keeps checking after everyone stops looking.

What this says about how I work.

I spent years delivering under MAS TRM, where nothing touches production without acceptance evidence, and a year at a global consultancy learning what auditors do to claims. Forward-deployed AI needs exactly that marriage: the pace of an engineer embedded at the client, and the paranoia of someone who knows the sample is not the season. So I pre-commit the criteria, make the trade-off visible enough that a business owner can genuinely choose, gate the signature on the evidence, and leave a harness behind that keeps asking the question after I have left the building. Fast is the job. Provable is the craft.

See it run, live.

Start at the vendor default and watch two criteria fail. Find the window, sign at your chosen point, then open week six and watch the same harness take the approval back. The evidence tab shows every error the accuracy number is made of.

Open the Go-Live Gate →Runs client-side over a synthetic labelled sample: no sign-in, nothing to install.