AI Hallucination Research › Briefings

Briefings Blog

The running blog from the RLB Specialist Panel delves into real-world scenarios where the compliance, legal, or AI lab team interacts with frontier AI models under specific regulations. The blogs are anonymised to remove client-specific details and include insights from the RLB team analysing the hallucinations experienced in AI models while working on these cases. For example, when a model returns a confident answer that contradicts the regulator's primary text, such as a fabricated staff letter, a wrong appendix, or an inverted scope, these issues are discussed here. Each blog explains one set of findings and what it would have meant for the team that would have acted on it, sans this research initiative. This blog is frequently updated, a few times a day.

263 briefings in the archive · Subscribe via Atom: /briefings/feed.xml (this blog) · /feed.xml (all RegLegBrief publications)
Audience colours: AI Labs Practitioner (profession) Sector × Department
Audience
Jur.
Regulator
Profession
Sector
Dept
Range
Sort
Per page
Showing 5 of 263 · page 34 of 53
Tuesday, 30 June 2026
Practitioner: Financial Advisers INT IMF

Financial Advisers: AI summaries of IMF Charges & Surcharge Reform (2024) may understate professional obligations

For Financial Advisers working with Review of Charges and the Surcharge Policy, Reform Proposals (October 2024): where Specialist-Panel-verified divergences between frontier AI summaries and the regulator's primary...

Financial advisers covering sovereign debt, multilateral lending, and IMF-program exposure are increasingly using AI to update client positioning notes on surcharge relief, generate sovereign-credit briefings on the the IMF October 2024 Surcharge Reform, and validate the headline 20-to-13 cohort figure against the IMF Board record before circulating to clients.

The RLB Specialist Panel put a set of practitioner-grade questions on the IMF October 2024 Surcharge Reform to two frontier AI models with web search active. Each question is prepared by the Panel based on the workflows that financial advisers actually use AI for under this reform, covering the pre-reform baseline of surcharge-paying members, the post-reform cohort projection through fiscal year 2026, and the immediate distributional impact of the 1 November 2024 effective date.

The Panel then binds every AI response to verbatim regulator-issued source text held as primary substrate, comparing the AI output line-by-line against the IMF Executive Board's published record. Only responses where the AI subject was demonstrably wrong against the verbatim regulator-issued source text are published; responses that were substantively correct, or that refused on calibration grounds, are retained internally and not surfaced.

On the IMF October 2024 Surcharge Reform, the AI subjects returned the same wrong cohort figure in the form of Numeric Drift, in the form of Inference Drift on one model and Outdated Retrieval on the other for financial advisers.

For financial advisers covering IMF-program-country exposure, multilateral debt, and sovereign credit, the cohort figure feeds directly into client positioning notes, peer-country comparison tables, and credit-relative-value frameworks. A client receiving advice anchored to a 19-country baseline rather than the Board's 20 receives advice that is factually off by one country on the cohort and off by one country on the relief count, with both errors traceable to the same off-by-one in the AI output.

The exposure is reputational: the client, or a peer adviser working the same trade, will check the figure against the IMF press release and identify the error on first review.

The published Specialist Panel findings, with model attribution, carry the following citation identifiers, each hyperlinked to the bound regulator-issued source text on the the IMF October 2024 Surcharge Reform regulation hub. The audit register surfaces these findings for financial advisers so that any AI-assisted figure entering a deliverable on the surcharge cohort, the FY2026 projection, or the per-country relief count can be re-validated against the IMF Executive Board record before the document is issued:

Sector: Investment Banking and Dept: Operations INT BIS-CPMI

Investment Banking Operations teams: documentation and reporting gaps possible from AI reading of PFMI (Principles for Financial Market Infrastructures)

For Investment Banking Operations teams working with Principles for Financial Market Infrastructures (PFMI): Specialist-Panel-verified findings on where AI summaries diverge from the regulator's text, and what that...

Operations teams at Investment Banking firms working on the CPMI-IOSCO Principles for Financial Market Infrastructures (PFMI, 2012) are increasingly relying on AI to draft CSP contract governance and escalation runbooks under an FMI's operating model, prepare incident-response playbooks for CSP performance failures, structure third-party performance monitoring against the FMI's mandate, and validate supervisor engagement protocols against the regulator-issued Annex F text.

The PFMI framework is the global standard for systemically important payment systems, central counterparties, and securities settlement infrastructures, and the document's structure makes it particularly amenable to AI summarisation: numbered Principles, numbered Key Considerations, and lettered annexes that the model can address by number.

That surface structure is also what makes the failure mode the RegLeg Brief Specialist Panel records here invisible at runtime: the document is regularly cited by Key Consideration number in board papers, disclosure-framework returns, and counterparty representations, which means a misattributed citation does not register as a substantive error in the draft, it registers as a competent regulatory paragraph that the reader will not check against the regulator's primary text unless something else prompts the verification.

Two frontier AI models tested by the RegLeg Brief Specialist Panel produced confidently wrong reconstructions of the PFMI's governance and oversight architecture under Principle 2 (governance) and Annex F (oversight expectations for critical service providers). The Panel records one finding in the class the team labels "Supervisor-Scope Inversion", in which the models stated a substantively plausible governance position and pinned it to a named Key Consideration that the published PFMI text does not support. The finding identifiers are RLB-H-INT-BIS-CPMI-IOSCO-PFMI-2012-Q011-Sonnet46.

For Operations teams at Investment Banking firms, the failure shape matters because the work product is CSP-relationship governance procedures, incident-response and escalation runbooks, third-party performance monitoring schedules, and supervisor-engagement protocols for outsourced services, all of which travel under the firm's name to a board, supervisor, counterparty, or public reviewer who can locate the cited Key Consideration and check it against the regulator's primary text.

Operations teams at investment banks owning the live operating model for CSP relationships under an FMI mandate are the population most exposed when AI output frames supervisor reach as ending at the FMI boundary, because the runbook will not anticipate a regulator-driven inquiry that contacts the CSP directly under Annex F.

The Panel documents the finding identifiers RLB-H-INT-BIS-CPMI-IOSCO-PFMI-2012-Q011-Sonnet46. The AI subjects under test were Claude Sonnet 4.6, each running with web search enabled, mirroring the workflow most practitioners run when they ask an assistant a Principle 2 or Annex F question. The verbatim regulator text is held as primary substrate (R2-REGULATION-d101a_PFMI_main_text.pdf). Each finding card sets out the exact strings the model produced, the verbatim regulator excerpt the model's output contradicts, and the failure-class label the RegLeg Brief Specialist Panel assigns.

The records are open-access; AI labs named in any finding have an unconditional right of reply, and the Specialist Panel will document any factual correction or contextual response alongside the original finding.

Sector: Investment Banking and Dept: Legal INT BIS-CPMI

Investment Banking Legal teams: documentation and reporting gaps possible from AI reading of PFMI (Principles for Financial Market Infrastructures)

For Investment Banking Legal teams working with Principles for Financial Market Infrastructures (PFMI): Specialist-Panel-verified findings on where AI summaries diverge from the regulator's text, and what that means...

Legal teams at Investment Banking firms working on the CPMI-IOSCO Principles for Financial Market Infrastructures (PFMI, 2012) are increasingly relying on AI to draft legal opinions on FMI governance obligations, structure third-party oversight clauses in CSP contracts under an FMI mandate, prepare counterparty representations on PFMI compliance, and validate committee-architecture language in board terms of reference against the regulator's Key Consideration text.

The PFMI framework is the global standard for systemically important payment systems, central counterparties, and securities settlement infrastructures, and the document's structure makes it particularly amenable to AI summarisation: numbered Principles, numbered Key Considerations, and lettered annexes that the model can address by number.

That surface structure is also what makes the failure mode the RegLeg Brief Specialist Panel records here invisible at runtime: the document is regularly cited by Key Consideration number in board papers, disclosure-framework returns, and counterparty representations, which means a misattributed citation does not register as a substantive error in the draft, it registers as a competent regulatory paragraph that the reader will not check against the regulator's primary text unless something else prompts the verification.

Two frontier AI models tested by the RegLeg Brief Specialist Panel produced confidently wrong reconstructions of the PFMI's governance and oversight architecture under Principle 2 (governance) and Annex F (oversight expectations for critical service providers). The Panel records two findings in the class the team labels "Source-Credit Fabrication and Supervisor-Scope Inversion", in which the models stated a substantively plausible governance position and pinned it to a named Key Consideration that the published PFMI text does not support. The finding identifiers are RLB-H-INT-BIS-CPMI-IOSCO-PFMI-2012-Q011-Sonnet46, RLB-H-INT-BIS-CPMI-IOSCO-PFMI-2012-Q022-Opus47.

For Legal teams at Investment Banking firms, the failure shape matters because the work product is legal opinions on FMI board obligations, CSP contract clauses on supervisor-engagement scope, counterparty representations on PFMI compliance, and drafted board terms of reference, all of which travel under the firm's name to a board, supervisor, counterparty, or public reviewer who can locate the cited Key Consideration and check it against the regulator's primary text.

Legal teams at investment banks signing off on opinions, contract structures, or board documentation are the population most exposed when an AI output assigns a non-existent obligation to a named Key Consideration, because the opinion carries the firm's name to the counterparty or supervisor who will check the citation against the PFMI primary text.

The Panel documents the finding identifiers RLB-H-INT-BIS-CPMI-IOSCO-PFMI-2012-Q011-Sonnet46; RLB-H-INT-BIS-CPMI-IOSCO-PFMI-2012-Q022-Opus47. The AI subjects under test were Claude Opus 4.7 and Claude Sonnet 4.6, each running with web search enabled, mirroring the workflow most practitioners run when they ask an assistant a Principle 2 or Annex F question. The verbatim regulator text is held as primary substrate (R2-REGULATION-d101a_PFMI_main_text.pdf). Each finding card sets out the exact strings the model produced, the verbatim regulator excerpt the model's output contradicts, and the failure-class label the RegLeg Brief Specialist Panel assigns.

The records are open-access; AI labs named in any finding have an unconditional right of reply, and the Specialist Panel will document any factual correction or contextual response alongside the original finding.

Sector: Payment Institutions and Dept: Legal INT BIS-CPMI

Payment Institutions Legal teams: documentation and reporting gaps possible from AI reading of PFMI (Principles for Financial Market Infrastructures)

For Payment Institutions Legal teams working with Principles for Financial Market Infrastructures (PFMI): Specialist-Panel-verified findings on where AI summaries diverge from the regulator's text, and what that...

Legal teams at Payment Institutions firms working on the CPMI-IOSCO Principles for Financial Market Infrastructures (PFMI, 2012) are increasingly relying on AI to draft legal opinions on FMI third-party oversight obligations, structure CSP-mandate contracts and contractual flow-down clauses, prepare counterparty representations on PFMI Annex F compliance, and validate supervisor-engagement scope language against the regulator's Annex F text.

The PFMI framework is the global standard for systemically important payment systems, central counterparties, and securities settlement infrastructures, and the document's structure makes it particularly amenable to AI summarisation: numbered Principles, numbered Key Considerations, and lettered annexes that the model can address by number.

That surface structure is also what makes the failure mode the RegLeg Brief Specialist Panel records here invisible at runtime: the document is regularly cited by Key Consideration number in board papers, disclosure-framework returns, and counterparty representations, which means a misattributed citation does not register as a substantive error in the draft, it registers as a competent regulatory paragraph that the reader will not check against the regulator's primary text unless something else prompts the verification.

Two frontier AI models tested by the RegLeg Brief Specialist Panel produced confidently wrong reconstructions of the PFMI's governance and oversight architecture under Principle 2 (governance) and Annex F (oversight expectations for critical service providers). The Panel records one finding in the class the team labels "Supervisor-Scope Inversion", in which the models stated a substantively plausible governance position and pinned it to a named Key Consideration that the published PFMI text does not support. The finding identifiers are RLB-H-INT-BIS-CPMI-IOSCO-PFMI-2012-Q011-Sonnet46.

For Legal teams at Payment Institutions firms, the failure shape matters because the work product is legal opinions on FMI third-party oversight, CSP-mandate contract structures, counterparty representations on PFMI Annex F compliance, and supervisor-engagement scope memoranda, all of which travel under the firm's name to a board, supervisor, counterparty, or public reviewer who can locate the cited Key Consideration and check it against the regulator's primary text.

Legal teams at payment institutions signing off on opinions or contract structures for CSP mandates are the population most exposed when AI output documents the supervisory relationship as purely contractual and FMI-internal, because the opinion carries the firm's name to a counterparty or supervisor who can locate Annex F and see the parallel regulator-to-CSP oversight channel the text contemplates.

The Panel documents the finding identifiers RLB-H-INT-BIS-CPMI-IOSCO-PFMI-2012-Q011-Sonnet46. The AI subjects under test were Claude Sonnet 4.6, each running with web search enabled, mirroring the workflow most practitioners run when they ask an assistant a Principle 2 or Annex F question. The verbatim regulator text is held as primary substrate (R2-REGULATION-d101a_PFMI_main_text.pdf). Each finding card sets out the exact strings the model produced, the verbatim regulator excerpt the model's output contradicts, and the failure-class label the RegLeg Brief Specialist Panel assigns.

The records are open-access; AI labs named in any finding have an unconditional right of reply, and the Specialist Panel will document any factual correction or contextual response alongside the original finding.

Sector: Payment Institutions and Dept: Governance & Company Secretarial INT BIS-CPMI

Payment Institutions Governance & Company Secretarial teams: documentation and reporting gaps possible from AI reading of PFMI (Principles for Financial Market Infrastructures)

For Payment Institutions Governance & Company Secretarial teams working with Principles for Financial Market Infrastructures (PFMI): Specialist-Panel-verified findings on where AI summaries diverge from the...

Governance & Company Secretarial teams at Payment Institutions firms working on the CPMI-IOSCO Principles for Financial Market Infrastructures (PFMI, 2012) are increasingly relying on AI to draft board charters and committee terms of reference under PFMI Principle 2, prepare board papers describing risk-management frameworks, complete PFMI disclosure-framework templates for the board's governance section, and validate committee-mandate language against the regulator-issued Key Considerations.

The PFMI framework is the global standard for systemically important payment systems, central counterparties, and securities settlement infrastructures, and the document's structure makes it particularly amenable to AI summarisation: numbered Principles, numbered Key Considerations, and lettered annexes that the model can address by number.

That surface structure is also what makes the failure mode the RegLeg Brief Specialist Panel records here invisible at runtime: the document is regularly cited by Key Consideration number in board papers, disclosure-framework returns, and counterparty representations, which means a misattributed citation does not register as a substantive error in the draft, it registers as a competent regulatory paragraph that the reader will not check against the regulator's primary text unless something else prompts the verification.

Two frontier AI models tested by the RegLeg Brief Specialist Panel produced confidently wrong reconstructions of the PFMI's governance and oversight architecture under Principle 2 (governance) and Annex F (oversight expectations for critical service providers). The Panel records one finding in the class the team labels "Source-Credit Misattribution", in which the models stated a substantively plausible governance position and pinned it to a named Key Consideration that the published PFMI text does not support. The finding identifiers are RLB-H-INT-BIS-CPMI-IOSCO-PFMI-2012-Q022-Sonnet46.

For Governance & Company Secretarial teams at Payment Institutions firms, the failure shape matters because the work product is board charters, risk-committee terms of reference, governance policy manuals, PFMI disclosure-framework responses, and committee-mandate submissions to the board, all of which travel under the firm's name to a board, supervisor, counterparty, or public reviewer who can locate the cited Key Consideration and check it against the regulator's primary text.

Governance and Company Secretarial teams at payment institutions responsible for committee architecture and the governance section of the disclosure-framework return are the population most exposed when AI output ties a committee recommendation to the wrong Key Consideration, because the return is reviewed against the PFMI primary text and the misattribution is visible to any reviewer who locates the cited Key Consideration.

The Panel documents the finding identifiers RLB-H-INT-BIS-CPMI-IOSCO-PFMI-2012-Q022-Sonnet46. The AI subjects under test were Claude Sonnet 4.6, each running with web search enabled, mirroring the workflow most practitioners run when they ask an assistant a Principle 2 or Annex F question. The verbatim regulator text is held as primary substrate (R2-REGULATION-d101a_PFMI_main_text.pdf). Each finding card sets out the exact strings the model produced, the verbatim regulator excerpt the model's output contradicts, and the failure-class label the RegLeg Brief Specialist Panel assigns.

The records are open-access; AI labs named in any finding have an unconditional right of reply, and the Specialist Panel will document any factual correction or contextual response alongside the original finding.

↑ Back to top