AI Hallucination Research › Briefings

Briefings Blog

The running blog from the RLB Specialist Panel delves into real-world scenarios where the compliance, legal, or AI lab team interacts with frontier AI models under specific regulations. The blogs are anonymised to remove client-specific details and include insights from the RLB team analysing the hallucinations experienced in AI models while working on these cases. For example, when a model returns a confident answer that contradicts the regulator's primary text, such as a fabricated staff letter, a wrong appendix, or an inverted scope, these issues are discussed here. Each blog explains one set of findings and what it would have meant for the team that would have acted on it, sans this research initiative. This blog is frequently updated, a few times a day.

263 briefings in the archive · Subscribe via Atom: /briefings/feed.xml (this blog) · /feed.xml (all RegLegBrief publications)
Audience colours: AI Labs Practitioner (profession) Sector × Department
Audience
Jur.
Regulator
Profession
Sector
Dept
Range
Sort
Per page
Showing 5 of 263 · page 2 of 53
Saturday, 25 July 2026
Practitioner: Lawyers INT BIS-CPMI

Lawyers: AI summaries of PFMI (Principles for Financial Market Infrastructures) may understate professional obligations

For Lawyers working with Principles for Financial Market Infrastructures (PFMI): where Specialist-Panel-verified divergences between frontier AI summaries and the regulator's primary source can affect client work,...

Lawyers working on the CPMI-IOSCO Principles for Financial Market Infrastructures (PFMI, 2012) are increasingly relying on AI to draft 2-page board memos on FMI governance obligations, generate client-facing summaries of PFMI Principle 2 Key Considerations, prepare partner-level briefings on disclosure-framework responses, and validate committee-mandate language against the regulator's published Key Considerations. The PFMI framework is the global standard for systemically important payment systems, central counterparties, and securities settlement infrastructures, and the document's structure makes it particularly amenable to AI summarisation: numbered Principles, numbered Key Considerations, and lettered annexes that the model can address by number.

That surface structure is also what makes the failure mode the RegLeg Brief Specialist Panel records here invisible at runtime: the document is regularly cited by Key Consideration number in board papers, disclosure-framework returns, and counterparty representations, which means a misattributed citation does not register as a substantive error in the draft, it registers as a competent regulatory paragraph that the reader will not check against the regulator's primary text unless something else prompts the verification.

Two frontier AI models tested by the RegLeg Brief Specialist Panel produced confidently wrong reconstructions of the PFMI's governance and oversight architecture under Principle 2 (governance) and Annex F (oversight expectations for critical service providers). The Panel records one finding in the class the team labels "Source-Credit Misattribution", in which the models stated a substantively plausible governance position and pinned it to a named Key Consideration that the published PFMI text does not support. The finding identifiers are RLB-H-INT-BIS-CPMI-IOSCO-PFMI-2012-Q022-Sonnet46.

For Lawyers, the failure shape matters because the work product is client memos on FMI board obligations, legal opinions on PFMI self-assessment responses, drafted board terms of reference, and counterparty representations on governance compliance, all of which travel under the firm's name to a board, supervisor, counterparty, or public reviewer who can locate the cited Key Consideration and check it against the regulator's primary text.

Lawyers who paste AI output into a legal opinion or a client-facing compliance memo are the population most exposed to a fabricated Key Consideration citation that the reader will check against the PFMI primary text the moment a counterparty or supervisor disputes the position.

The Panel documents the finding identifiers RLB-H-INT-BIS-CPMI-IOSCO-PFMI-2012-Q022-Sonnet46. The AI subjects under test were Claude Sonnet 4.6, each running with web search enabled, mirroring the workflow most practitioners run when they ask an assistant a Principle 2 or Annex F question. The verbatim regulator text is held as primary substrate (R2-REGULATION-d101a_PFMI_main_text.pdf). Each finding card sets out the exact strings the model produced, the verbatim regulator excerpt the model's output contradicts, and the failure-class label the RegLeg Brief Specialist Panel assigns.

The records are open-access; AI labs named in any finding have an unconditional right of reply, and the Specialist Panel will document any factual correction or contextual response alongside the original finding.

Practitioner: Lawyers INT BIS-CPMI

Lawyers: AI summaries of CPMI-IOSCO Cyber Resilience for FMIs (2016) may understate professional obligations

For Lawyers working with Guidance on Cyber Resilience for Financial Market Infrastructures (CPMI-IOSCO 2016): where Specialist-Panel-verified divergences between frontier AI summaries and the regulator's primary...

Lawyers advising on cyber resilience for financial market infrastructures and the CPMI-IOSCO 2016 Cyber Guidance are increasingly using AI to draft client memos, validate threshold language, and prepare partner-level briefings on the global guidance and its post-2016 evolution. In practice, AI is used to draft client memos on the CPMI-IOSCO 2016 Cyber Guidance, validate cyber-programme citations against the regulator text, generate partner-level briefings on how the guidance is referenced by national supervisors, and prepare counsel-to-board commentary on FMI cyber-resilience standards.

That workflow places the regulator-issued text of the 2016 guidance, its 2018-2020 derivative standards, and its current operative status at the centre of every AI-generated deliverable for lawyers.

Two frontier AI models tested by the RegLeg Brief Specialist Panel produced confident, citable reconstructions of the CPMI-IOSCO 2016 Cyber Guidance (June 2016) that the regulator-issued primary text directly contradicts across nine findings spanning four failure classes: Source-Credit Fabrication (an asserted NIST Cybersecurity Framework citation that the 2016 guidance does not contain), Misattribution (the slogan 'secure the periphery, protect the core' located inside CPMI-IOSCO 2016 guidance or its 2018 wholesale-payments paper rather than the actual 2018 speech source), Anachronistic Cross-Reference (the 2016 guidance asserted as definitionally aligned with the November 2018 FSB Cyber Lexicon and the October 2020 FSB Effective Practices that postdate it), and Outdated Standing Claim (the 2016 guidance presented as the unchanged operative standard when CPMI-IOSCO has issued a May 2026 consultative document under active revision).

Questions are prepared by the RLB Specialist Panel based on real practical AI usage in the workflows lawyers use AI for. The Panel binds each AI finding to verbatim regulator-issued source text held as primary substrate.

For lawyers advising FMI operators, supervisors, and FMI participant banks, the failure pattern is operationally consequential. A client memorandum that recites an explicit NIST CSF citation that the 2016 guidance does not contain misstates the regulatory foundation. A counsel-to-board briefing that records the 2016 guidance as the unchanged operative standard, when CPMI-IOSCO has issued a May 2026 consultative document under active revision, embeds a falsifiable status claim into a regulated deliverable.

The audit's nine findings are documented with immutable RLB Citation IDs. Representative entries include RLB-H-INT-BIS-CPMI-IOSCO-CYBER-RESILIENCE-FMI-2016-Q008-Opus47, RLB-H-INT-BIS-CPMI-IOSCO-CYBER-RESILIENCE-FMI-2016-Q008-Sonnet46, RLB-H-INT-BIS-CPMI-IOSCO-CYBER-RESILIENCE-FMI-2016-Q014-Opus47, RLB-H-INT-BIS-CPMI-IOSCO-CYBER-RESILIENCE-FMI-2016-Q014-Sonnet46, RLB-H-INT-BIS-CPMI-IOSCO-CYBER-RESILIENCE-FMI-2016-Q019-Sonnet46, RLB-H-INT-BIS-CPMI-IOSCO-CYBER-RESILIENCE-FMI-2016-Q020-Opus47, RLB-H-INT-BIS-CPMI-IOSCO-CYBER-RESILIENCE-FMI-2016-Q020-Sonnet46, RLB-H-INT-BIS-CPMI-IOSCO-CYBER-RESILIENCE-FMI-2016-Q022-Opus47, and RLB-H-INT-BIS-CPMI-IOSCO-CYBER-RESILIENCE-FMI-2016-Q022-Sonnet46. The full audit is documented at the CPMI-IOSCO 2016 Cyber Resilience Guidance hub on RegLegBrief.com.

Practitioner: Company Secretaries INT BIS-CPMI

Company Secretaries: AI summaries of PFMI (Principles for Financial Market Infrastructures) may understate professional obligations

For Company Secretaries working with Principles for Financial Market Infrastructures (PFMI): where Specialist-Panel-verified divergences between frontier AI summaries and the regulator's primary source can affect...

Company Secretaries working on the CPMI-IOSCO Principles for Financial Market Infrastructures (PFMI, 2012) are increasingly relying on AI to draft board and risk-committee terms of reference, prepare papers for the board's oversight of critical service providers, validate committee mandates against the PFMI's published Key Considerations, and assemble governance disclosures for the FMI's annual disclosure-framework return. The PFMI framework is the global standard for systemically important payment systems, central counterparties, and securities settlement infrastructures, and the document's structure makes it particularly amenable to AI summarisation: numbered Principles, numbered Key Considerations, and lettered annexes that the model can address by number.

That surface structure is also what makes the failure mode the RegLeg Brief Specialist Panel records here invisible at runtime: the document is regularly cited by Key Consideration number in board papers, disclosure-framework returns, and counterparty representations, which means a misattributed citation does not register as a substantive error in the draft, it registers as a competent regulatory paragraph that the reader will not check against the regulator's primary text unless something else prompts the verification.

Two frontier AI models tested by the RegLeg Brief Specialist Panel produced confidently wrong reconstructions of the PFMI's governance and oversight architecture under Principle 2 (governance) and Annex F (oversight expectations for critical service providers). The Panel records two findings in the class the team labels "Source-Credit Fabrication and Supervisor-Scope Inversion", in which the models stated a substantively plausible governance position and pinned it to a named Key Consideration that the published PFMI text does not support. The finding identifiers are RLB-H-INT-BIS-CPMI-IOSCO-PFMI-2012-Q011-Sonnet46, RLB-H-INT-BIS-CPMI-IOSCO-PFMI-2012-Q022-Opus47.

For Company Secretaries, the failure shape matters because the work product is board charters, risk-committee terms of reference, board information papers on third-party oversight, and PFMI disclosure-framework responses, all of which travel under the firm's name to a board, supervisor, counterparty, or public reviewer who can locate the cited Key Consideration and check it against the regulator's primary text. Company Secretaries who route AI-drafted board and committee documentation into the FMI's governance pack are the population most exposed when the model misnumbers a Key Consideration or fabricates a committee-architecture mandate that the PFMI text does not contain.

The Panel documents the finding identifiers RLB-H-INT-BIS-CPMI-IOSCO-PFMI-2012-Q011-Sonnet46; RLB-H-INT-BIS-CPMI-IOSCO-PFMI-2012-Q022-Opus47. The AI subjects under test were Claude Opus 4.7 and Claude Sonnet 4.6, each running with web search enabled, mirroring the workflow most practitioners run when they ask an assistant a Principle 2 or Annex F question. The verbatim regulator text is held as primary substrate (R2-REGULATION-d101a_PFMI_main_text.pdf). Each finding card sets out the exact strings the model produced, the verbatim regulator excerpt the model's output contradicts, and the failure-class label the RegLeg Brief Specialist Panel assigns.

The records are open-access; AI labs named in any finding have an unconditional right of reply, and the Specialist Panel will document any factual correction or contextual response alongside the original finding.

Practitioner: Company Secretaries INT BIS-CPMI

Company Secretaries: AI summaries of CPMI-IOSCO Cyber Resilience for FMIs (2016) may understate professional obligations

For Company Secretaries working with Guidance on Cyber Resilience for Financial Market Infrastructures (CPMI-IOSCO 2016): where Specialist-Panel-verified divergences between frontier AI summaries and the regulator's...

Company secretaries supporting FMI boards and corporate boards exposed to CPMI-IOSCO 2016 cyber-resilience expectations are increasingly using AI to draft board papers, prepare director-induction material, and maintain regulator horizon-scanning packs on the cyber-resilience framework. In practice, AI is used to draft board papers on the FMI cyber-resilience programme, populate director-induction material on the CPMI-IOSCO 2016 framework, prepare audit-committee briefings on cyber-supervisory expectations, and maintain the regulator horizon-scanning pack covering CPMI-IOSCO, FSB, and national supervisor publications.

That workflow places the regulator-issued text of the 2016 guidance, its 2018-2020 derivative standards, and its current operative status at the centre of every AI-generated deliverable for company secretaries.

Two frontier AI models tested by the RegLeg Brief Specialist Panel produced confident, citable reconstructions of the CPMI-IOSCO 2016 Cyber Guidance (June 2016) that the regulator-issued primary text directly contradicts across nine findings spanning four failure classes: Source-Credit Fabrication (an asserted NIST Cybersecurity Framework citation that the 2016 guidance does not contain), Misattribution (the slogan 'secure the periphery, protect the core' located inside CPMI-IOSCO 2016 guidance or its 2018 wholesale-payments paper rather than the actual 2018 speech source), Anachronistic Cross-Reference (the 2016 guidance asserted as definitionally aligned with the November 2018 FSB Cyber Lexicon and the October 2020 FSB Effective Practices that postdate it), and Outdated Standing Claim (the 2016 guidance presented as the unchanged operative standard when CPMI-IOSCO has issued a May 2026 consultative document under active revision).

Questions are prepared by the RLB Specialist Panel based on real practical AI usage in the workflows company secretaries use AI for. The Panel binds each AI finding to verbatim regulator-issued source text held as primary substrate.

For company secretaries supporting the board on the FMI cyber programme, the failure pattern is operationally consequential. A board paper that recites an explicit NIST CSF alignment of the 2016 guidance lands inside the paper as a regulator-grounded foundation claim. An induction pack that records the 2016 guidance and the November 2018 FSB Cyber Lexicon as definitionally aligned papers over a two-year vocabulary gap. A horizon-scanning pack that records the 2016 guidance as standing without active revision misses the May 2026 CPMI-IOSCO consultative document.

The audit's nine findings are documented with immutable RLB Citation IDs. Representative entries include RLB-H-INT-BIS-CPMI-IOSCO-CYBER-RESILIENCE-FMI-2016-Q008-Opus47, RLB-H-INT-BIS-CPMI-IOSCO-CYBER-RESILIENCE-FMI-2016-Q008-Sonnet46, RLB-H-INT-BIS-CPMI-IOSCO-CYBER-RESILIENCE-FMI-2016-Q014-Opus47, RLB-H-INT-BIS-CPMI-IOSCO-CYBER-RESILIENCE-FMI-2016-Q014-Sonnet46, RLB-H-INT-BIS-CPMI-IOSCO-CYBER-RESILIENCE-FMI-2016-Q019-Sonnet46, RLB-H-INT-BIS-CPMI-IOSCO-CYBER-RESILIENCE-FMI-2016-Q020-Opus47, RLB-H-INT-BIS-CPMI-IOSCO-CYBER-RESILIENCE-FMI-2016-Q020-Sonnet46, RLB-H-INT-BIS-CPMI-IOSCO-CYBER-RESILIENCE-FMI-2016-Q022-Opus47, and RLB-H-INT-BIS-CPMI-IOSCO-CYBER-RESILIENCE-FMI-2016-Q022-Sonnet46. The full audit is documented at the CPMI-IOSCO 2016 Cyber Resilience Guidance hub on RegLegBrief.com.

Sector: Management Consulting and Dept: Finance INT IMF

Management Consulting Finance teams: documentation and reporting gaps possible from AI reading of IMF Precautionary Balances 2026

For Management Consulting Finance teams working with Review of the Adequacy of the Fund's Precautionary Balances (2026): Specialist-Panel-verified findings on where AI summaries diverge from the regulator's text, and...

Frontier AI models tested against the International Monetary Fund's March 2026 Review of the Adequacy of the Fund's Precautionary Balances produced six confident, citable answers that the regulator's own primary text directly contradicts, an evaluation by the RLB Specialist Panel has found.

For Management Consulting Finance teams who use AI tools on Fund financial-governance and Fund-strength tracking matters, the failures concentrate on the specific numerical and lexicon parameters that are most commonly carried into working deliverables: the precautionary balances minimum floor, the FY2024 surcharge-payer baseline, the half-year PB level reported in the Q2FY26 Quarterly Financial Report, the strength of the Board's early-review signal, and the named geopolitical theatre the Board flagged as a source of intensifying downside risk.

The most material of the six findings concerns the minimum floor for precautionary balances. The IMF's own text on the March 2026 Review records that Directors generally agreed to retain the current floor at SDR 20 billion. The frontier AI model under test committed to a floor of SDR 15 billion across multiple deliverable registers, including a board briefing memo for an EM finance ministry client and a historical-trajectory section for a campaign report by a development-finance NGO.

The SDR 5 billion divergence is a verifiable parameter that supervisors, counterparties, and internal sign-off reviewers will check against the source; if it enters a deliverable for Management Consulting Finance teams use it will surface under review.

A second cluster of findings concerns the October 2024 charges and surcharge reform. The IMF Press Release records that the number of countries subject to surcharges in fiscal year 2026 is expected to fall from 20 to 13. The AI's policy-brief draft inflated the FY2024 baseline to 22, producing a 22-to-13 trajectory that diverges from the regulator's 20-to-13. A third finding concerns the IMF Board lexicon: the regulator's text records that 'a few Directors' saw merit in considering an early review of charges and the surcharge policy.

The AI's legal-and-policy advisory elevated the position to 'a number of Directors', changing the strength of the signal a sovereign-debt practitioner would read out for multi-year debt-service planning. A fourth finding adds Ukraine to the Board's named geopolitical theatre on intensifying downside risk; the regulator's text names only the Middle East.

A fifth and sixth finding record divergences on the half-year precautionary balances level reported in the Q2FY26 Quarterly Financial Report and on the pre-March-2024 floor value in a campaign-report historical trajectory. The Q2FY26 Schedule 2 records the October 31, 2025 PB level at SDR 26,782 million; the AI committed to approximately SDR 26.5 billion. Every finding in this audit is bound to verbatim primary source text recorded by the International Monetary Fund. The RLB Specialist Panel offers International Monetary Fund and any other named entity a permanent right of reply on every finding.

For Management Consulting Finance teams, the operational signal is that AI-assisted research on the March 2026 PB Review and the related October 2024 surcharge reform cannot be relied on for the floor value, the FY surcharge-payer baseline, the Board-lexicon strength of a Board signal, the named geopolitical theatre, or the half-year PB level, without verification against the IMF Press Release, the Q2FY26 Quarterly Financial Report, and Press Release 24/376 directly.

↑ Back to top