Every pain point below is a data problem wearing a compliance costume, and every credible AI answer has the same shape: AI at the edges, determinism at the core. Probabilistic components read documents, classify transactions and rank chases as confidence-scored proposals behind a human gate. The calculation engine, factor library, boundary register and hash-chained ledger never contain a model, so a disclosed figure re-derives by arithmetic. Anyone can put a language model inside the number. The valuable move is putting it around the number and keeping the number provable. This page is the evidence behind the case study: an adversarially verified research pass over EFRAG's cost-benefit study of first-wave CSRD preparers, IFAC's assurance state-of-play data, IAASB implementation guidance and peer-reviewed data-quality literature, three verification votes per claim.
Evidence labels used throughout: [verified 3-0] survived all three adversarial votes against primary sources · [verified 2-1] majority · [vendor] published by a party selling the solution, never load-bearing · [historical] period-limited.
EFRAG's study of first-wave CSRD preparers ranks the top three burdens as interpreting the standards, collecting value-chain data, and the double materiality assessment; reducing the number of data points ranks only fourth. [ranking verified 2-1; components 3-0] Simply understanding the standards consumed 15% of internal implementation cost. 89% of respondents rated the effort high or very high, none low; the median cost was about EUR 500,000 internal plus EUR 300,000 external, roughly half recurring annually. [verified 3-0; 43 self-selected Wave 1 respondents, commissioned alongside a simplification consultation: read as what respondents reported]
Reporting platforms assume you already know what the standard requires. They give you fields to fill, not answers to "does this apply to us, and what does this paragraph mean for a leased building?"
Retrieval-augmented answering over the standards themselves, with paragraph-level citation and abstention when the text does not cover the question. This is the single most AI-addressable line item in the EFRAG dataset, and an interpretation layer aimed at the largest soft cost in the regime is a product decision.
EFRAG ranks value-chain data collection the second-largest burden. [verified 3-0] Incompleteness is the normal state: firms disclosing a Scope 3 category breakdown reported an average of only 3.75 of the 15 GHG Protocol categories across 2010-2019, rising from 1.7 to 4.7 over the period. [verified 2-1, historical] For FY2023, 82% of GHG-reporting large listed companies disclosed any Scope 3 at all, with purchased goods (Category 1) covered by just 65%. [verified 3-0] For real estate the boundary problem is structural: CapitaLand's assured FY2025 disclosure puts Scope 3 at 65% of its footprint, with downstream leased assets alone at 63% of material Scope 3. [company-published, assured FY2025 disclosure] You cannot meter what you do not control.
The general ledger already contains most of Scope 1 and much of Scope 3, captured and audited for another purpose. And the IAASB has already set the auditability bar for mining it: ISSA 5000 implementation guidance treats a carbon platform that interfaces with the general ledger and auto-assigns emission factors as a control point, where the entity must satisfy itself that the factor database is appropriate and the assignment process works. [verified 3-0 on source content; the AI extension is the author's inference]
EFRAG ranks it the third-largest burden, at 20% of internal implementation cost, equal to preparing the disclosures themselves. [verified 3-0] It is labour-intensive because it is judgement-intensive: identify candidate topics, score impact and financial materiality, engage stakeholders, set thresholds, and document why each line landed where it did. Platforms provide matrices to fill in; they do not do the identification, the scoring rationale or the evidence trail, which is where the time goes.
First-pass topic identification from peers, sector standards, regulatory text and news; drafting scoring rationale for human editing; assembling the evidence pack. ISSA 5000 expects documented judgements including close calls, so a system that captures who scored what, on what scale, with what evidence, converts the most expensive soft activity into a reusable asset.
The standard-setter concedes it outright: ISSA 5000 implementation guidance states that entity process and control over sustainability information "may often be less than fully developed", particularly for first-time preparers. [verified 3-0] The market reflects it. In FY2023, 82% of engagements were limited assurance and only about 8% reasonable, with nearly two-thirds of that reasonable work covering GHG metrics only, so roughly 92% of assured sustainability data has never been tested to audit-equivalent depth. Non-audit providers signed 45% of assurance reports, yet only 38% of them applied ISAE 3000, 24% the IESBA ethics code and 41% quality-management standards, against 98% on both for audit firms. [verified 3-0; FY2023, superseded edition noted; the non-audit share has since receded to roughly 41%] And the cost hides off the fee line: internal assurance effort runs at roughly 25% of external fees for large preparers, with some interviewed companies at 1:1, friction attributed to ambiguity in the standards. [verified 2-1]
Continuous evidence assembly at capture time rather than at fieldwork; drafting working papers, the basis of preparation and judgement memos from the audit trail rather than from memory; pre-fieldwork self-testing that recomputes every disclosure and triages exceptions before the auditor asks.
Pairwise ESG rating correlations across six major agencies run at only 0.38 to 0.71, average 0.54: divergence, not error, but divergence the capital markets price. [Berg et al., verified 3-0] For Scope 3, two major data vendors agree on only 68% of data points within 1% error, while model-estimated data matches neither on any data point, with a trimmed mean absolute percentage error of 111%. [verified 3-0] Modelled Scope 3 estimates are wrong by roughly the size of the answer. It is a direct argument for measured, FM-native data over estimated data, and a direct warning about AI systems that generate estimates rather than capture evidence.
Its job here is shrinking the estimated share: extraction from documents, meter and sensor capture, supplier primary-data collection, and anomaly detection that flags when a series stops behaving like measurement. Where estimates remain unavoidable, the contribution is disclosure hygiene: label the method, show the primary versus secondary share, and never let an estimate be promoted to a measurement silently.
The load-bearing negative finding deserves its own section: no verified evidence survived showing that AI is credibly fixing any of these pain points. Every adoption statistic and effectiveness claim traced back to a vendor selling the product. That includes widely circulated survey figures from carbon-measurement vendors: genuinely published, but validating the publisher's own product thesis, so this pass's verifiers refused them as independent evidence. They are labelled [vendor] wherever they appear and are load-bearing nowhere.
Three implications. The market is loud and unproven, so the differentiated position belongs to whoever first demonstrates AI under the auditability bar the IAASB has already written. The bar is knowable: the worked example on GL-to-emissions platforms tells you exactly what an assurer will ask. And the honest claim is available: not "our AI reduces reporting effort by X%", which nobody can substantiate, but "here is where the model is allowed to stand, what it proposed, who approved it, and the arithmetic that produced the number." Provability is the product.
Primary: EFRAG Cost-Benefit Analysis on the Draft Amended ESRS (December 2025; 170 survey responses, 32 interviews; CSRD figures above from the 43-respondent Wave 1 subset); IFAC/AICPA & CIMA State of Play in Sustainability Assurance (FY2023 data, since superseded); IAASB ISSA 5000 implementation guidance; Berg et al. on ESG rating divergence; peer-reviewed work on Scope 3 data quality and category coverage; CapitaLand's FY2025 assured sustainability disclosure for the real-estate boundary figures. Known weaknesses, stated rather than buried: EFRAG's sample is self-selected and the study was commissioned alongside a simplification agenda; the Scope 3 category study ends in 2019, before CSRD and IFRS S2; IFAC's FY2023 figures are superseded and will be refreshed; practitioner-voice sources did not survive verification, so the qualitative texture of complaint is missing here. Vendor-published figures are labelled as such throughout and carry no load.