WTWilly Tai
Case study · live demoSustainability · ESG Assurance

Your sustainability report reads beautifully. Then someone is paid to check it.

From FY2029, Singapore-listed companies face limited assurance over Scope 1 and 2 emissions, and reports written for readers start being read by practitioners. This bench runs a twelve-requirement climate checklist (the IFRS S2 four-pillar structure: governance, strategy, risk management, metrics and targets) against a synthetic sustainability report, with two disciplines the market's AI-powered "ESG checkers" skip: every verdict must cite an exact quote that the bench mechanically verifies exists, or abstain; and an assurance lens re-scores the same report with narrative paying zero, because that is how it will be read when checking becomes a paid engagement. Runs live in your browser.

01 / The pain, and why the naive approach fails

Reports are graded on prose today and re-performance tomorrow, and most were written for today.

A decade of voluntary sustainability reporting selected for narrative: commitments, intentions, "in due course". Assurance selects for something else entirely: a number, its boundary, its method, and a data trail a practitioner can walk. The gap between those two readings is invisible until someone flips the lens, which is precisely what FY2029 does to every listed issuer's Scope 1 and 2. The naive tool here is an LLM that "reviews your ESG report" and produces confident summaries with invented page references; the failure mode is well documented and fatal for this use case, because a checklist opinion that cites a paragraph which does not exist is worse than no opinion. So the bench inverts the design: the model's role (in a live engagement) is to propose candidate passages, and the bench only accepts a verdict whose quote string-matches the report text. No match, no citation; no citation, an abstention that says exactly what is missing.

02 / One worked example

Flip the lens and watch the strategy section evaporate.

Under the reader's lens the report scores respectably: governance 75 percent, metrics a full 100. Flip to the assurer's lens and the same twelve verdicts re-price: metrics fall to 60 percent as "we will consider setting formal targets as methodologies mature" stops counting, and strategy falls from 50 to 11, because a capital-planning linkage asserted in one sentence cannot be re-performed. What survives the flip is exactly what should: Scope 1 at 1,842 tCO2e with boundary and factors stated, Scope 2 with its 68.2 GWh and 41 metered facilities, and a data-quality paragraph honest enough to disclose that 8 percent of area is estimated. One requirement, scenario analysis, gets a clean abstention: nothing in the report resembles it, and the bench does not fill gaps.

PillarReader's lensAssurer's lens · FY2029
Governance75%67%
Strategy50%11%
Risk management75%67%
Metrics & targets100%60%
The mechanical citation check is the differentiator. Eleven verdicts cite quotes; all eleven are verified against the report text on every load, and a citation that failed verification would downgrade itself to absent, visibly. That is the discipline that makes an AI-assisted version of this bench trustworthy: the model proposes, the string match decides, and the abstention is a first-class verdict.
03 / Numbers & honest trade-offs

What it does, and what it does and doesn't prove.

12
requirements, 4 pillars
Paraphrasing the IFRS S2 structure.
11 / 1
citations / abstentions
Every quote mechanically verified in the text.
100 → 60
metrics, at the lens flip
What narrative was worth all along.
FY2029
the deadline that matters
Scope 1 and 2 limited assurance, listed issuers.

Where this stands, honestly. The report is synthetic, written to contain the exact mix assurers meet; the requirement texts are paraphrases of the four-pillar structure, not the standard's words, and twelve requirements is a bench, not the full checklist a real engagement runs. The scoring weights (narrative zero, partial a third under the assurance lens) are a design argument rather than a standard's rule. What transfers exactly is the mechanism and the message: cite or abstain, verified mechanically, and read yourself the hard way before someone is paid to. This bench pairs with my delivery record deliberately: the metering coverage paragraph it rewards is the kind of sensor-to-report trail my smart-building programmes actually build. The proposition in one line: I read your report the way FY2029 will, three years early, and hand you the gap list while it is still cheap.

What this says about how I work.

The two halves of my repositioning meet on this page: the sustainability numbers come from buildings I know how to instrument, and the checking discipline comes from the audit training I am reactivating. I built the bench cite-or-abstain because that is the only honest architecture for AI near assurance: a model that must show a verifiable quote cannot hallucinate a compliance opinion, and a bench that abstains loudly is worth more than one that fills gaps politely. The lens flip is the client conversation in one click: not "is your report good", but "which sentences survive when checking becomes a fee".

See it run, live.

Run the checklist under the reader's lens, flip to the assurer's lens and watch the pillars re-price, then open the report tab and read the five paragraphs nothing cites, because an assurer reads the gaps first.

Open the bench →Runs client-side over a synthetic report: no sign-in, nothing to install.