AI Hallucination Research › Briefings

Briefings Blog

The running blog from the RLB Specialist Panel delves into real-world scenarios where the compliance, legal, or AI lab team interacts with frontier AI models under specific regulations. The blogs are anonymised to remove client-specific details and include insights from the RLB team analysing the hallucinations experienced in AI models while working on these cases. For example, when a model returns a confident answer that contradicts the regulator's primary text, such as a fabricated staff letter, a wrong appendix, or an inverted scope, these issues are discussed here. Each blog explains one set of findings and what it would have meant for the team that would have acted on it, sans this research initiative. This blog is frequently updated, a few times a day.

263 briefings in the archive · Subscribe via Atom: /briefings/feed.xml (this blog) · /feed.xml (all RegLegBrief publications)
Audience colours: AI Labs Practitioner (profession) Sector × Department
Audience
Jur.
Regulator
Profession
Sector
Dept
Range
Sort
Per page
Showing 5 of 263 · page 1 of 53
Sunday, 26 July 2026
Practitioner: Lawyers INT BIS-CPMI

Lawyers: AI summaries of CPMI ISO 20022 Harmonisation (2026 update) may understate professional obligations

For Lawyers working with Harmonised ISO 20022 Data Requirements for Enhancing Cross-Border Payments - Updated Report: where Specialist-Panel-verified divergences between frontier AI summaries and the regulator's...

Lawyers advising on the CPMI Harmonised ISO 20022 Data Requirements (Updated Report) are increasingly using AI to draft client memos on Fedwire postal address compliance, validate threshold language in correspondent banking opinions, and prepare partner-level briefings on the harmonised ISO 20022 governance lineage. The same tools support first-pass advice on cross-border payment obligations under the updated CPMI data model.

Two frontier AI models tested by the RLB Specialist Panel on the workflows lawyers use to support advice on the CPMI Harmonised ISO 20022 Data Requirements (Updated Report) produced three discrete hallucinations bound to regulator-issued source text. The Panel records two distinct failure classes, Numeric Drift and Schema Over-Specification across the set. Questions were prepared by the Specialist Panel based on real practical AI usage in the workflows lawyers use AI for, and each finding is bound to verbatim regulator-issued source text held as primary substrate.

For Lawyers, each hallucination has a direct operational consequence in the regulatory opinion, partner-level memo, or correspondent banking advice. The Panel's testing surfaces ISO 20022 adoption rate conflation (RTGS vs faster payments), Fedwire hybrid postal address schema over-specification, and ISO 20022 adoption rate conflation (RTGS vs faster payments). Where these errors flow into a deliverable, the exposure is PI exposure, client correction, and discoverable error in opinion drafts that propagate to multiple counterparties.

The pattern is uniform across the set: the AI returns a confident, sourced-looking answer that conflicts in a load-bearing specific with the regulator's verbatim text, and the error survives a first-pass review precisely because the surface form is plausible. The Panel records each hallucination with the regulator's primary substrate held as the anchor, so the corrective text is available alongside the failure.

The Specialist Panel records the citation IDs as follows: RLB-H-INT-BIS-CPMI-ISO-20022-HARMONISATION-UPDATED-2026-Q006-Opus47 (Claude Opus 4.7 (web search on), Numeric Drift); RLB-H-INT-BIS-CPMI-ISO-20022-HARMONISATION-UPDATED-2026-Q010-Opus47 (Claude Opus 4.7 (web search on), Schema Over-Specification); RLB-H-INT-BIS-CPMI-ISO-20022-HARMONISATION-UPDATED-2026-Q006-Sonnet46 (Claude Sonnet 4.6 (web search on), Numeric Drift). Each citation links to the verbatim regulator-issued source text, the tested AI question, and the recorded AI response, so the Panel's assessment is traceable end to end. For lawyers, the citation IDs operate as a reference index: when an AI answer in the working draft matches a known Panel finding, the cited regulator text is already available as the corrective anchor.

The full per-finding analysis cards, including the audience-specific impact statement, sit on the cell's detail surface for sign-off use.

Practitioner: Company Secretaries INT BIS-CPMI

Company Secretaries: AI summaries of CPMI ISO 20022 Harmonisation (2026 update) may understate professional obligations

For Company Secretaries working with Harmonised ISO 20022 Data Requirements for Enhancing Cross-Border Payments - Updated Report: where Specialist-Panel-verified divergences between frontier AI summaries and the...

Company Secretaries supporting boards exposed to the CPMI Harmonised ISO 20022 Data Requirements (Updated Report) are increasingly using AI to draft board papers on the governance lineage of payment-system standards, generate director-onboarding notes on cross-border payments policy obligations, and prepare regulatory horizon-scanning entries on CPMI workstreams. The same tools validate the attribution of standard-setting authority in committee minutes.

Two frontier AI models tested by the RLB Specialist Panel on the workflows company secretaries use to support advice on the CPMI Harmonised ISO 20022 Data Requirements (Updated Report) produced one discrete hallucination bound to regulator-issued source text. The Panel records a single recurring failure class: Source-Credit Fabrication across the set. Questions were prepared by the Specialist Panel based on real practical AI usage in the workflows company secretaries use AI for, and each finding is bound to verbatim regulator-issued source text held as primary substrate.

For Company Secretaries, each hallucination has a direct operational consequence in the board paper, director-onboarding note, or horizon-scanning entry. The Panel's testing surfaces CPMI working-group chair misattribution. Where these errors flow into a deliverable, the exposure is board-pack error that compounds across multiple committee cycles before it surfaces in a director question or supervisory review. The pattern is uniform across the set: the AI returns a confident, sourced-looking answer that conflicts in a load-bearing specific with the regulator's verbatim text, and the error survives a first-pass review precisely because the surface form is plausible.

The Panel records each hallucination with the regulator's primary substrate held as the anchor, so the corrective text is available alongside the failure.

The Specialist Panel records the citation IDs as follows: RLB-H-INT-BIS-CPMI-ISO-20022-HARMONISATION-UPDATED-2026-Q004-Sonnet46 (Claude Sonnet 4.6 (web search on), Source-Credit Fabrication). Each citation links to the verbatim regulator-issued source text, the tested AI question, and the recorded AI response, so the Panel's assessment is traceable end to end. For company secretaries, the citation IDs operate as a reference index: when an AI answer in the working draft matches a known Panel finding, the cited regulator text is already available as the corrective anchor. The full per-finding analysis cards, including the audience-specific impact statement, sit on the cell's detail surface for sign-off use.

Practitioner: Accountants (CA/PA) INT BIS-CPMI

Accountants (CA/PA): AI summaries of CPMI ISO 20022 Harmonisation (2026 update) may understate professional obligations

For Accountants (CA/PA) working with Harmonised ISO 20022 Data Requirements for Enhancing Cross-Border Payments - Updated Report: where Specialist-Panel-verified divergences between frontier AI summaries and the...

Accountants advising on the CPMI Harmonised ISO 20022 Data Requirements (Updated Report) are increasingly using AI to model the ROI of ISO 20022 migration for client business cases, populate audit working papers with regulator-sourced inquiry-rate benchmarks, and validate adoption-rate citations in advisory deliverables. The same tools draft sections of client board memos that depend on official CPMI or FSB quantitative anchors.

Two frontier AI models tested by the RLB Specialist Panel on the workflows accountants use to support advice on the CPMI Harmonised ISO 20022 Data Requirements (Updated Report) produced three discrete hallucinations bound to regulator-issued source text. The Panel records two distinct failure classes, False-Negative Retrieval and Numeric Drift across the set. Questions were prepared by the Specialist Panel based on real practical AI usage in the workflows accountants use AI for, and each finding is bound to verbatim regulator-issued source text held as primary substrate.

For Accountants (CA/PA), each hallucination has a direct operational consequence in the audit working paper, advisory memo, or client board paper. The Panel's testing surfaces ISO 20022 adoption rate conflation (RTGS vs faster payments), missing inquiry-rate and resolution-time benchmarks, and ISO 20022 adoption rate conflation (RTGS vs faster payments). Where these errors flow into a deliverable, the exposure is misstated peer benchmark, weakened investment case, and credibility loss when clients verify the underlying source.

The pattern is uniform across the set: the AI returns a confident, sourced-looking answer that conflicts in a load-bearing specific with the regulator's verbatim text, and the error survives a first-pass review precisely because the surface form is plausible. The Panel records each hallucination with the regulator's primary substrate held as the anchor, so the corrective text is available alongside the failure.

The Specialist Panel records the citation IDs as follows: RLB-H-INT-BIS-CPMI-ISO-20022-HARMONISATION-UPDATED-2026-Q006-Opus47 (Claude Opus 4.7 (web search on), Numeric Drift); RLB-H-INT-BIS-CPMI-ISO-20022-HARMONISATION-UPDATED-2026-Q007-Sonnet46 (Claude Sonnet 4.6 (web search on), False-Negative Retrieval); RLB-H-INT-BIS-CPMI-ISO-20022-HARMONISATION-UPDATED-2026-Q006-Sonnet46 (Claude Sonnet 4.6 (web search on), Numeric Drift). Each citation links to the verbatim regulator-issued source text, the tested AI question, and the recorded AI response, so the Panel's assessment is traceable end to end. For accountants (ca/pa), the citation IDs operate as a reference index: when an AI answer in the working draft matches a known Panel finding, the cited regulator text is already available as the corrective anchor.

The full per-finding analysis cards, including the audience-specific impact statement, sit on the cell's detail surface for sign-off use.

Practitioner: Professional Engineers INT UNTC

Professional Engineers: AI summaries of BBNJ Agreement may understate professional obligations

For Professional Engineers working with BBNJ High Seas Biodiversity Agreement: where Specialist-Panel-verified divergences between frontier AI summaries and the regulator's primary source can affect client work,...

Professional engineers scoping projects that may touch areas beyond national jurisdiction are increasingly using AI to draft environmental impact assessment scoping documents, generate technical briefings for design teams on screening thresholds, and validate which provision of the BBNJ Agreement governs a particular obligation before submitting deliverables to clients or regulators.

The RLB Specialist Panel put a set of practitioner-grade questions on the BBNJ Agreement to two frontier AI models with web search active. Each question is prepared by the Panel based on the workflows that professional engineers actually use AI for under this treaty, covering the screening threshold for environmental impact assessments under Part IV, the temporal scope of the marine genetic resources and digital sequence information regime under Part II, the benefit-sharing duty for digital sequence information, and the non-undermining duty constraining Conference of the Parties decisions on area-based management tools under Part III.

The Panel then binds every AI response to verbatim regulator-issued source text held as primary substrate, comparing the AI output line-by-line against the deposited treaty text. Only responses where the AI subject was demonstrably wrong against the verbatim regulator-issued source text are published; responses that were substantively correct, or that refused on calibration grounds, are retained internally and not surfaced. On the BBNJ Agreement, the AI subjects returned a single hallucinated answer in the form of Source-Credit Misattribution for professional engineers.

For professional engineers scoping projects that may engage areas beyond national jurisdiction, citation accuracy in environmental impact assessment scoping documents is load-bearing. A scoping document that pins the screening obligation to the wrong article of the BBNJ Agreement will be challenged on first review by a regulator, a peer reviewer, or the client's own legal team. The substantive screening test the AI paraphrased may be the right test, but the citation will need to be reworked, and the engineer issuing the document carries professional responsibility for the accuracy of the regulatory reference framing the work.

The published Specialist Panel findings, with model attribution, carry the following citation identifiers, each hyperlinked to the bound regulator-issued source text on the BBNJ Agreement regulation hub. The audit register surfaces these findings for professional engineers so that any AI-assisted treaty citation, paraphrase, or rule-statement entering a deliverable can be re-validated against the deposited treaty text before the document is issued:

Practitioner: Lawyers INT UNTC

Lawyers: AI summaries of BBNJ Agreement may understate professional obligations

For Lawyers working with BBNJ High Seas Biodiversity Agreement: where Specialist-Panel-verified divergences between frontier AI summaries and the regulator's primary source can affect client work, professional...

Lawyers advising on the BBNJ Agreement are increasingly using AI to draft 2-page client memos on benefit-sharing exposure, generate partner-level briefings on environmental impact assessment thresholds for high-seas activities, and validate treaty-citation language against the deposited Agreement text before issuing legal opinions.

The RLB Specialist Panel put a set of practitioner-grade questions on the BBNJ Agreement to two frontier AI models with web search active. Each question is prepared by the Panel based on the workflows that lawyers actually use AI for under this treaty, covering the screening threshold for environmental impact assessments under Part IV, the temporal scope of the marine genetic resources and digital sequence information regime under Part II, the benefit-sharing duty for digital sequence information, and the non-undermining duty constraining Conference of the Parties decisions on area-based management tools under Part III.

The Panel then binds every AI response to verbatim regulator-issued source text held as primary substrate, comparing the AI output line-by-line against the deposited treaty text. Only responses where the AI subject was demonstrably wrong against the verbatim regulator-issued source text are published; responses that were substantively correct, or that refused on calibration grounds, are retained internally and not surfaced. On the BBNJ Agreement, the AI subjects returned four hallucinated answers in the form of Inverted-Position Hallucination together with Source-Credit Misattribution for lawyers.

For lawyers issuing legal opinions, memoranda, and transactional documents that engage the BBNJ Agreement, treaty-citation accuracy is load-bearing: a counterparty, opposing counsel, or regulatory reviewer who can identify a citation error on first reading of the document calls the entire piece of advice into question.

An AI-drafted memo that points at the wrong article on a screening threshold, mis-states the direction of a benefit-sharing rule, or mis-locates the constraint on Conference of the Parties authority leaves the lawyer exposed to professional liability, the firm exposed to reputational risk, and the client exposed to commercial loss from a position structured on the wrong rule.

The published Specialist Panel findings, with model attribution, carry the following citation identifiers, each hyperlinked to the bound regulator-issued source text on the BBNJ Agreement regulation hub. The audit register surfaces these findings for lawyers so that any AI-assisted treaty citation, paraphrase, or rule-statement entering a deliverable can be re-validated against the deposited treaty text before the document is issued:

↑ Back to top