WTWilly Tai
Ledger Notes · No. 1 · CommentaryEvidence base

Five pain points in ESG reporting, and where AI actually helps.

Every pain point below is a data problem wearing a compliance costume, and every credible AI answer has the same shape: AI at the edges, determinism at the core. Probabilistic components read documents, classify transactions and rank chases as confidence-scored proposals behind a human gate. The calculation engine, factor library, boundary register and hash-chained ledger never contain a model, so a disclosed figure re-derives by arithmetic. Anyone can put a language model inside the number. The valuable move is putting it around the number and keeping the number provable. This page is the evidence behind the case study: an adversarially verified research pass over EFRAG's cost-benefit study of first-wave CSRD preparers, IFAC's assurance state-of-play data, IAASB implementation guidance and peer-reviewed data-quality literature, three verification votes per claim.

Evidence labels used throughout: [verified 3-0] survived all three adversarial votes against primary sources · [verified 2-1] majority · [vendor] published by a party selling the solution, never load-bearing · [historical] period-limited.

Two honesty notes before anything else.
1. The AI half of this question is unevidenced. Every claim offering AI adoption rates or effectiveness data came from vendor-published surveys and was refuted in adversarial verification, in several cases explicitly because the vendor sells the tool the survey validates. There is no credible independent evidence that AI is currently fixing any of these pain points. That is not a reason to avoid AI; it is the strategic opening, and the final section explains why.
2. Practitioner-voice research failed. Nothing from forums, professional-network threads or first-report retrospectives survived verification. Everything below comes from institutional surveys, standard-setters and academic work: the pain points are structurally evidenced but not emotionally documented.
Pain 1 / Interpreting the standards

The number one complaint is understanding what the standard means, not filling it in.

EFRAG's study of first-wave CSRD preparers ranks the top three burdens as interpreting the standards, collecting value-chain data, and the double materiality assessment; reducing the number of data points ranks only fourth. [ranking verified 2-1; components 3-0] Simply understanding the standards consumed 15% of internal implementation cost. 89% of respondents rated the effort high or very high, none low; the median cost was about EUR 500,000 internal plus EUR 300,000 external, roughly half recurring annually. [verified 3-0; 43 self-selected Wave 1 respondents, commissioned alongside a simplification consultation: read as what respondents reported]

Why software has not solved it

Reporting platforms assume you already know what the standard requires. They give you fields to fill, not answers to "does this apply to us, and what does this paragraph mean for a leased building?"

Where AI earns its place

Retrieval-augmented answering over the standards themselves, with paragraph-level citation and abstention when the text does not cover the question. This is the single most AI-addressable line item in the EFRAG dataset, and an interpretation layer aimed at the largest soft cost in the regime is a product decision.

Pain 2 / Data you do not own

The big numbers belong to other people, and the rest are scattered across bills, portals and ledgers.

EFRAG ranks value-chain data collection the second-largest burden. [verified 3-0] Incompleteness is the normal state: firms disclosing a Scope 3 category breakdown reported an average of only 3.75 of the 15 GHG Protocol categories across 2010-2019, rising from 1.7 to 4.7 over the period. [verified 2-1, historical] For FY2023, 82% of GHG-reporting large listed companies disclosed any Scope 3 at all, with purchased goods (Category 1) covered by just 65%. [verified 3-0] For real estate the boundary problem is structural: CapitaLand's assured FY2025 disclosure puts Scope 3 at 65% of its footprint, with downstream leased assets alone at 63% of material Scope 3. [company-published, assured FY2025 disclosure] You cannot meter what you do not control.

The general ledger already contains most of Scope 1 and much of Scope 3, captured and audited for another purpose. And the IAASB has already set the auditability bar for mining it: ISSA 5000 implementation guidance treats a carbon platform that interfaces with the general ledger and auto-assigns emission factors as a control point, where the entity must satisfy itself that the factor database is appropriate and the assignment process works. [verified 3-0 on source content; the AI extension is the author's inference]

The AI stack for this pain

  1. Document intelligence on bills, fuel delivery notes, service reports and waste tickets, with confidence scores and the source retained.
  2. Ledger-line classification to emission categories, learning from the chart of accounts and prior confirmations: the highest-leverage idea available, now with a known audit bar to clear.
  3. A learned source registry: the system remembers where each site's data came from, in what format, from whom, and fetches it next period. First cycle expensive, fifth cycle nearly free.
  4. An expected-data calendar with absence detection: the alert that matters is the number that never arrived.
  5. Lease and contract mining to find which agreements already grant submeter or data-sharing rights, converting a legal archive into a data strategy.
Guardrail. Extraction is a proposal, never a posting. Below a confidence threshold it queues for a human. AI-derived values are flagged in the lineage exactly like gap-filled estimates and never blended silently into measured data.
Pain 3 / Double materiality

The materiality assessment costs as much as preparing every disclosure, because it is judgement all the way down.

EFRAG ranks it the third-largest burden, at 20% of internal implementation cost, equal to preparing the disclosures themselves. [verified 3-0] It is labour-intensive because it is judgement-intensive: identify candidate topics, score impact and financial materiality, engage stakeholders, set thresholds, and document why each line landed where it did. Platforms provide matrices to fill in; they do not do the identification, the scoring rationale or the evidence trail, which is where the time goes.

Where AI earns its place

First-pass topic identification from peers, sector standards, regulatory text and news; drafting scoring rationale for human editing; assembling the evidence pack. ISSA 5000 expects documented judgements including close calls, so a system that captures who scored what, on what scale, with what evidence, converts the most expensive soft activity into a reusable asset.

Guardrail. The model proposes candidates and drafts rationale. Thresholds and final scoring are human decisions, recorded with an owner and a date. A materiality assessment authored by a model is exactly the artefact an assurer will attack first.
Pain 4 / The controls gap

The most under-priced pain in the market: the controls under the numbers do not exist yet.

The standard-setter concedes it outright: ISSA 5000 implementation guidance states that entity process and control over sustainability information "may often be less than fully developed", particularly for first-time preparers. [verified 3-0] The market reflects it. In FY2023, 82% of engagements were limited assurance and only about 8% reasonable, with nearly two-thirds of that reasonable work covering GHG metrics only, so roughly 92% of assured sustainability data has never been tested to audit-equivalent depth. Non-audit providers signed 45% of assurance reports, yet only 38% of them applied ISAE 3000, 24% the IESBA ethics code and 41% quality-management standards, against 98% on both for audit firms. [verified 3-0; FY2023, superseded edition noted; the non-audit share has since receded to roughly 41%] And the cost hides off the fee line: internal assurance effort runs at roughly 25% of external fees for large preparers, with some interviewed companies at 1:1, friction attributed to ambiguity in the standards. [verified 2-1]

Where AI earns its place

Continuous evidence assembly at capture time rather than at fieldwork; drafting working papers, the basis of preparation and judgement memos from the audit trail rather than from memory; pre-fieldwork self-testing that recomputes every disclosure and triages exceptions before the auditor asks.

Guardrail. Models draft narrative; numbers are transcluded from the engine and never typed by a model. Every sentence must resolve to a record. An unsourced sentence is a defect, not a style issue.
Pain 5 / Unreliable at source

The data disagrees with itself, and everyone who buys it can see that.

Pairwise ESG rating correlations across six major agencies run at only 0.38 to 0.71, average 0.54: divergence, not error, but divergence the capital markets price. [Berg et al., verified 3-0] For Scope 3, two major data vendors agree on only 68% of data points within 1% error, while model-estimated data matches neither on any data point, with a trimmed mean absolute percentage error of 111%. [verified 3-0] Modelled Scope 3 estimates are wrong by roughly the size of the answer. It is a direct argument for measured, FM-native data over estimated data, and a direct warning about AI systems that generate estimates rather than capture evidence.

Where AI earns its place

Its job here is shrinking the estimated share: extraction from documents, meter and sensor capture, supplier primary-data collection, and anomaly detection that flags when a series stops behaving like measurement. Where estimates remain unavoidable, the contribution is disclosure hygiene: label the method, show the primary versus secondary share, and never let an estimate be promoted to a measurement silently.

The opening

Nobody has proven AI works here, and that is the position.

The load-bearing negative finding deserves its own section: no verified evidence survived showing that AI is credibly fixing any of these pain points. Every adoption statistic and effectiveness claim traced back to a vendor selling the product. That includes widely circulated survey figures from carbon-measurement vendors: genuinely published, but validating the publisher's own product thesis, so this pass's verifiers refused them as independent evidence. They are labelled [vendor] wherever they appear and are load-bearing nowhere.

Three implications. The market is loud and unproven, so the differentiated position belongs to whoever first demonstrates AI under the auditability bar the IAASB has already written. The bar is knowable: the worked example on GL-to-emissions platforms tells you exactly what an assurer will ask. And the honest claim is available: not "our AI reduces reporting effort by X%", which nobody can substantiate, but "here is where the model is allowed to stand, what it proposed, who approved it, and the arithmetic that produced the number." Provability is the product.

Sources and known weaknesses

Primary: EFRAG Cost-Benefit Analysis on the Draft Amended ESRS (December 2025; 170 survey responses, 32 interviews; CSRD figures above from the 43-respondent Wave 1 subset); IFAC/AICPA & CIMA State of Play in Sustainability Assurance (FY2023 data, since superseded); IAASB ISSA 5000 implementation guidance; Berg et al. on ESG rating divergence; peer-reviewed work on Scope 3 data quality and category coverage; CapitaLand's FY2025 assured sustainability disclosure for the real-estate boundary figures. Known weaknesses, stated rather than buried: EFRAG's sample is self-selected and the study was commissioned alongside a simplification agenda; the Scope 3 category study ends in 2019, before CSRD and IFRS S2; IFAC's FY2023 figures are superseded and will be refreshed; practitioner-voice sources did not survive verification, so the qualitative texture of complaint is missing here. Vendor-published figures are labelled as such throughout and carry no load.

CHANGELOG
2026-08-08 · First published, from the working thesis of 4 August 2026.
Planned · Refresh IFAC figures from the 2019-2024 edition; re-attempt practitioner-voice evidence with named-source interviews.
Back to the case study →The platform these findings shaped is walked through, exhibit by exhibit, in the case study.