AI Hallucination Research › Briefings

Briefings Blog

The running blog from the RLB Specialist Panel delves into real-world scenarios where the compliance, legal, or AI lab team interacts with frontier AI models under specific regulations. The blogs are anonymised to remove client-specific details and include insights from the RLB team analysing the hallucinations experienced in AI models while working on these cases. For example, when a model returns a confident answer that contradicts the regulator's primary text, such as a fabricated staff letter, a wrong appendix, or an inverted scope, these issues are discussed here. Each blog explains one set of findings and what it would have meant for the team that would have acted on it, sans this research initiative. This blog is frequently updated, a few times a day.

263 briefings in the archive · Subscribe via Atom: /briefings/feed.xml (this blog) · /feed.xml (all RegLegBrief publications)
Audience colours: AI Labs Practitioner (profession) Sector × Department
Audience
Jur.
Regulator
Profession
Sector
Dept
Range
Sort
Per page
Showing 5 of 263 · page 43 of 53
Tuesday, 23 June 2026
Sector: Payment Institutions and Dept: Compliance INT BIS-CPMI

Payment Institutions Compliance teams: documentation and reporting gaps possible from AI reading of CPMI ISO 20022 Harmonisation (2026 update)

For Payment Institutions Compliance teams working with Harmonised ISO 20022 Data Requirements for Enhancing Cross-Border Payments - Updated Report: Specialist-Panel-verified findings on where AI summaries diverge...

Compliance teams at Payment Institutions operating under the CPMI Harmonised ISO 20022 Data Requirements (Updated Report) are increasingly using AI to draft regulatory horizon-scanning records on adoption progress, generate correspondent-network readiness assessments, and validate the postal-address mapping in the firm's ISO 20022 message structure. The same tools prepare supervisor-facing descriptions of ISO 20022 readiness.

Two frontier AI models tested by the RLB Specialist Panel on the workflows payment-institution compliance officers use to support advice on the CPMI Harmonised ISO 20022 Data Requirements (Updated Report) produced three discrete hallucinations bound to regulator-issued source text. The Panel records two distinct failure classes, Numeric Drift and Schema Over-Specification across the set. Questions were prepared by the Specialist Panel based on real practical AI usage in the workflows payment-institution compliance officers use AI for, and each finding is bound to verbatim regulator-issued source text held as primary substrate.

For Compliance teams at Payment Institutions, each hallucination has a direct operational consequence in the horizon-scanning record, network-readiness assessment, or supervisor-facing readiness description. The Panel's testing surfaces ISO 20022 adoption rate conflation (RTGS vs faster payments), Fedwire hybrid postal address schema over-specification, and ISO 20022 adoption rate conflation (RTGS vs faster payments). Where these errors flow into a deliverable, the exposure is skewed correspondent-network readiness picture, over-specified vendor due-diligence criteria, and a discoverable error in the firm's regulatory record.

The pattern is uniform across the set: the AI returns a confident, sourced-looking answer that conflicts in a load-bearing specific with the regulator's verbatim text, and the error survives a first-pass review precisely because the surface form is plausible. The Panel records each hallucination with the regulator's primary substrate held as the anchor, so the corrective text is available alongside the failure.

The Specialist Panel records the citation IDs as follows: RLB-H-INT-BIS-CPMI-ISO-20022-HARMONISATION-UPDATED-2026-Q006-Opus47 (Claude Opus 4.7 (web search on), Numeric Drift); RLB-H-INT-BIS-CPMI-ISO-20022-HARMONISATION-UPDATED-2026-Q010-Opus47 (Claude Opus 4.7 (web search on), Schema Over-Specification); RLB-H-INT-BIS-CPMI-ISO-20022-HARMONISATION-UPDATED-2026-Q006-Sonnet46 (Claude Sonnet 4.6 (web search on), Numeric Drift). Each citation links to the verbatim regulator-issued source text, the tested AI question, and the recorded AI response, so the Panel's assessment is traceable end to end.

For compliance teams at payment institutions, the citation IDs operate as a reference index: when an AI answer in the working draft matches a known Panel finding, the cited regulator text is already available as the corrective anchor. The full per-finding analysis cards, including the audience-specific impact statement, sit on the cell's detail surface for sign-off use.

Monday, 22 June 2026
Sector: Retail Banking and Dept: Compliance INT BIS-CPMI

Retail Banking Compliance teams: documentation and reporting gaps possible from AI reading of CPMI ISO 20022 Harmonisation (2026 update)

For Retail Banking Compliance teams working with Harmonised ISO 20022 Data Requirements for Enhancing Cross-Border Payments - Updated Report: Specialist-Panel-verified findings on where AI summaries diverge from the...

Compliance teams at Retail Banking firms operating under the CPMI Harmonised ISO 20022 Data Requirements (Updated Report) are increasingly using AI to update correspondent-banking onboarding checklists against the CPMI data model, generate horizon-scanning entries for the payments business line, and prepare NED briefings on the regulator's adoption-rate posture. The same tools validate citation accuracy in compliance attestations and supervisory exchanges.

Two frontier AI models tested by the RLB Specialist Panel on the workflows retail-banking compliance officers use to support advice on the CPMI Harmonised ISO 20022 Data Requirements (Updated Report) produced three discrete hallucinations bound to regulator-issued source text. The Panel records two distinct failure classes, Numeric Drift and Source-Credit Fabrication across the set. Questions were prepared by the Specialist Panel based on real practical AI usage in the workflows retail-banking compliance officers use AI for, and each finding is bound to verbatim regulator-issued source text held as primary substrate.

For Compliance teams at Retail Banking firms, each hallucination has a direct operational consequence in the compliance attestation, NED briefing, or horizon-scanning entry. The Panel's testing surfaces CPMI working-group chair misattribution, ISO 20022 adoption rate conflation (RTGS vs faster payments), and ISO 20022 adoption rate conflation (RTGS vs faster payments). Where these errors flow into a deliverable, the exposure is misstated peer-group benchmark in NED packs and a discoverable factual error in supervisory exchanges.

The pattern is uniform across the set: the AI returns a confident, sourced-looking answer that conflicts in a load-bearing specific with the regulator's verbatim text, and the error survives a first-pass review precisely because the surface form is plausible. The Panel records each hallucination with the regulator's primary substrate held as the anchor, so the corrective text is available alongside the failure.

The Specialist Panel records the citation IDs as follows: RLB-H-INT-BIS-CPMI-ISO-20022-HARMONISATION-UPDATED-2026-Q004-Sonnet46 (Claude Sonnet 4.6 (web search on), Source-Credit Fabrication); RLB-H-INT-BIS-CPMI-ISO-20022-HARMONISATION-UPDATED-2026-Q006-Opus47 (Claude Opus 4.7 (web search on), Numeric Drift); RLB-H-INT-BIS-CPMI-ISO-20022-HARMONISATION-UPDATED-2026-Q006-Sonnet46 (Claude Sonnet 4.6 (web search on), Numeric Drift). Each citation links to the verbatim regulator-issued source text, the tested AI question, and the recorded AI response, so the Panel's assessment is traceable end to end.

For compliance teams at retail banking firms, the citation IDs operate as a reference index: when an AI answer in the working draft matches a known Panel finding, the cited regulator text is already available as the corrective anchor. The full per-finding analysis cards, including the audience-specific impact statement, sit on the cell's detail surface for sign-off use.

Sector: Payment Institutions and Dept: Product & Business Development INT BIS-CPMI

Payment Institutions Product & Business Development teams: documentation and reporting gaps possible from AI reading of CPMI ISO 20022 Harmonisation (2026 update)

For Payment Institutions Product & Business Development teams working with Harmonised ISO 20022 Data Requirements for Enhancing Cross-Border Payments - Updated Report: Specialist-Panel-verified findings on where AI...

Product & Business Development teams at Payment Institutions building cross-border payment propositions under the CPMI Harmonised ISO 20022 Data Requirements (Updated Report) are increasingly using AI to build corridor strategy and investor decks, draft business cases for enrichment investment, and populate partner pitch materials with regulator-sourced inquiry-rate benchmarks. The same tools validate competitive-positioning claims with CPMI adoption data.

Two frontier AI models tested by the RLB Specialist Panel on the workflows payment-institution product and business-development teams use to support advice on the CPMI Harmonised ISO 20022 Data Requirements (Updated Report) produced three discrete hallucinations bound to regulator-issued source text. The Panel records two distinct failure classes, False-Negative Retrieval and Numeric Drift across the set. Questions were prepared by the Specialist Panel based on real practical AI usage in the workflows payment-institution product and business-development teams use AI for, and each finding is bound to verbatim regulator-issued source text held as primary substrate.

For Product & Business Development teams at Payment Institutions, each hallucination has a direct operational consequence in the corridor strategy, investor deck, or business case for enrichment. The Panel's testing surfaces ISO 20022 adoption rate conflation (RTGS vs faster payments), missing inquiry-rate and resolution-time benchmarks, and ISO 20022 adoption rate conflation (RTGS vs faster payments). Where these errors flow into a deliverable, the exposure is reputational damage in partner conversations, an investment case with the wrong source attribution for the headline efficiency metric, and a product roadmap built on a false market assumption.

The pattern is uniform across the set: the AI returns a confident, sourced-looking answer that conflicts in a load-bearing specific with the regulator's verbatim text, and the error survives a first-pass review precisely because the surface form is plausible. The Panel records each hallucination with the regulator's primary substrate held as the anchor, so the corrective text is available alongside the failure.

The Specialist Panel records the citation IDs as follows: RLB-H-INT-BIS-CPMI-ISO-20022-HARMONISATION-UPDATED-2026-Q006-Opus47 (Claude Opus 4.7 (web search on), Numeric Drift); RLB-H-INT-BIS-CPMI-ISO-20022-HARMONISATION-UPDATED-2026-Q007-Sonnet46 (Claude Sonnet 4.6 (web search on), False-Negative Retrieval); RLB-H-INT-BIS-CPMI-ISO-20022-HARMONISATION-UPDATED-2026-Q006-Sonnet46 (Claude Sonnet 4.6 (web search on), Numeric Drift). Each citation links to the verbatim regulator-issued source text, the tested AI question, and the recorded AI response, so the Panel's assessment is traceable end to end.

For product & business development teams at payment institutions, the citation IDs operate as a reference index: when an AI answer in the working draft matches a known Panel finding, the cited regulator text is already available as the corrective anchor. The full per-finding analysis cards, including the audience-specific impact statement, sit on the cell's detail surface for sign-off use.

Sector: Corporate Banking and Dept: Compliance INT BIS-CPMI

Corporate Banking Compliance teams: documentation and reporting gaps possible from AI reading of CPMI ISO 20022 Harmonisation (2026 update)

For Corporate Banking Compliance teams working with Harmonised ISO 20022 Data Requirements for Enhancing Cross-Border Payments - Updated Report: Specialist-Panel-verified findings on where AI summaries diverge from...

Compliance teams at Corporate Banking firms responsible for cross-border payments under the CPMI Harmonised ISO 20022 Data Requirements (Updated Report) are increasingly using AI to draft compliance attestations on Fedwire payment-format configuration, populate internal control narratives, and generate audit walkthrough documentation on ISO 20022 mapping. The same tools verify peer-benchmark statistics in supervisory communications.

Two frontier AI models tested by the RLB Specialist Panel on the workflows corporate-banking compliance officers use to support advice on the CPMI Harmonised ISO 20022 Data Requirements (Updated Report) produced three discrete hallucinations bound to regulator-issued source text. The Panel records two distinct failure classes, Numeric Drift and Schema Over-Specification across the set. Questions were prepared by the Specialist Panel based on real practical AI usage in the workflows corporate-banking compliance officers use AI for, and each finding is bound to verbatim regulator-issued source text held as primary substrate.

For Compliance teams at Corporate Banking firms, each hallucination has a direct operational consequence in the compliance attestation, control narrative, or supervisory communication. The Panel's testing surfaces ISO 20022 adoption rate conflation (RTGS vs faster payments), Fedwire hybrid postal address schema over-specification, and ISO 20022 adoption rate conflation (RTGS vs faster payments). Where these errors flow into a deliverable, the exposure is vendor specification built on an over-specified format, AML data-quality gaps, and a discoverable factual error in supervisory submissions.

The pattern is uniform across the set: the AI returns a confident, sourced-looking answer that conflicts in a load-bearing specific with the regulator's verbatim text, and the error survives a first-pass review precisely because the surface form is plausible. The Panel records each hallucination with the regulator's primary substrate held as the anchor, so the corrective text is available alongside the failure.

The Specialist Panel records the citation IDs as follows: RLB-H-INT-BIS-CPMI-ISO-20022-HARMONISATION-UPDATED-2026-Q006-Opus47 (Claude Opus 4.7 (web search on), Numeric Drift); RLB-H-INT-BIS-CPMI-ISO-20022-HARMONISATION-UPDATED-2026-Q010-Opus47 (Claude Opus 4.7 (web search on), Schema Over-Specification); RLB-H-INT-BIS-CPMI-ISO-20022-HARMONISATION-UPDATED-2026-Q006-Sonnet46 (Claude Sonnet 4.6 (web search on), Numeric Drift). Each citation links to the verbatim regulator-issued source text, the tested AI question, and the recorded AI response, so the Panel's assessment is traceable end to end.

For compliance teams at corporate banking firms, the citation IDs operate as a reference index: when an AI answer in the working draft matches a known Panel finding, the cited regulator text is already available as the corrective anchor. The full per-finding analysis cards, including the audience-specific impact statement, sit on the cell's detail surface for sign-off use.

Sector: Hedge Funds and Dept: Risk US CFTC

Hedge Funds Risk teams: documentation and reporting gaps possible from AI reading of CFTC Regulation 1.25 (Customer Funds Investments)

For Hedge Funds Risk teams working with Amendments to Regulation 1.25, Permissible Investments of Customer Funds by Futures Commission Merchants and Derivatives Clearing Organizations: Specialist-Panel-verified...

Risk teams at hedge fund managers monitoring FCM clearing-broker concentration exposure under Regulation 1.25 are increasingly using frontier AI assistants to produce FCM clearing-broker concentration-exposure dashboards, validate DWAM scenario assumptions against the clearing-broker's reported carve-out set, prepare counterparty risk briefings on the 2024 amendments, and to surface practical readings of the 2024 amendment package issued by the Commodity Futures Trading Commission (CFTC) on permissible investments of customer segregated funds under Regulation 1.25.

The amendments restate the 50 per cent concentration ceiling for government money market funds and qualified Treasury ETFs, the 24-month portfolio dollar-weighted average maturity (DWAM) standard and its carve-out set, and the separate March 31, 2025 compliance anchor for the Segregation Investment Detail Report (SIDR) and customer risk disclosure statement updates. Across this question set the model outputs that risk teams at hedge fund managers would carry into a FCM clearing-broker exposure dashboards departed from the regulator's verbatim text on each of the three operative axes.

Two frontier AI models tested by the RegLeg Brief (RLB) Specialist Panel reproduced the same failure shape across the audited question set on the CFTC's 2024 amendments to Regulation 1.25 (permissible investments of customer segregated funds by futures commission merchants and derivatives clearing organizations). The Panel calls the pattern Threshold-Trigger Elision and Carve-Out Inversion. The frontier AI models dropped the asset-size and management-company-size triggers that activate the 50 per cent concentration ceiling, swapped U.S. Treasury repurchase agreements into the DWAM exclusion set in place of the regulator's actual three carved-out classes, returned a no-DWAM-standard answer for direct U.S.

Treasury obligations where the 24-month portfolio standard governs by default, and drifted from the March 31, 2025 SIDR compliance anchor into a generic "roughly six months to a year after the effective date" formulation. The Panel records the failure class as inference_drift across the five audited findings, each bound to verbatim regulator-issued primary substrate held by the Panel.

For risk teams at hedge fund managers the operational consequence is direct. A clearing-broker exposure dashboard built on a uniform 50 per cent ceiling would understate the rule's scope and the clearing-broker's actual concentration posture. A DWAM scenario assumption that accepts U.S. Treasury repos as a carved-out class would model the wrong portfolio decomposition. A counterparty risk briefing anchored to a relative SIDR range would misadvise the risk committee on the regulator's March 31, 2025 anchor.

The failure surfaces in workflows the audience already uses AI for, the model output reads as a fluent reconstruction of the amended rule, and validation only happens if the reader independently knew the dual-trigger structure of the 50 per cent ceiling, the three-class DWAM carve-out, and the March 31, 2025 SIDR anchor. None of these are properties the audience can recover at runtime from the AI output alone.

The five findings are published with immutable RLB Citation IDs and bound to verbatim Commodity Futures Trading Commission source text: RLB-H-US-CFTC-FCM-DCO-CUSTOMER-FUNDS-INVESTMENTS-REG-1-25-2024-Q001-Opus47, RLB-H-US-CFTC-FCM-DCO-CUSTOMER-FUNDS-INVESTMENTS-REG-1-25-2024-Q001-Sonnet46, RLB-H-US-CFTC-FCM-DCO-CUSTOMER-FUNDS-INVESTMENTS-REG-1-25-2024-Q002-Opus47, RLB-H-US-CFTC-FCM-DCO-CUSTOMER-FUNDS-INVESTMENTS-REG-1-25-2024-Q002-Sonnet46, RLB-H-US-CFTC-FCM-DCO-CUSTOMER-FUNDS-INVESTMENTS-REG-1-25-2024-Q004-Opus47. The full audit on Regulation 1.25 is on the Regulation 1.25 (2024 amendments) hub on RegLegBrief.com.

↑ Back to top