AI Hallucination Research › Briefings

Briefings Blog

The running blog from the RLB Specialist Panel delves into real-world scenarios where the compliance, legal, or AI lab team interacts with frontier AI models under specific regulations. The blogs are anonymised to remove client-specific details and include insights from the RLB team analysing the hallucinations experienced in AI models while working on these cases. For example, when a model returns a confident answer that contradicts the regulator's primary text, such as a fabricated staff letter, a wrong appendix, or an inverted scope, these issues are discussed here. Each blog explains one set of findings and what it would have meant for the team that would have acted on it, sans this research initiative. This blog is frequently updated, a few times a day.

263 briefings in the archive · Subscribe via Atom: /briefings/feed.xml (this blog) · /feed.xml (all RegLegBrief publications)
Audience colours: AI Labs Practitioner (profession) Sector × Department
Audience
Jur.
Regulator
Profession
Sector
Dept
Range
Sort
Per page
Showing 3 of 263 · page 53 of 53
Saturday, 13 June 2026
AI Labs INT IMF

Specialist Panel: Frontier AI models misread IMF Charges & Surcharge Reform (2024)

RegLegBrief's Specialist Panel finds frontier AI models with web search enabled diverge from the regulator's verbatim text of Review of Charges and the Surcharge Policy, Reform Proposals (October 2024). Findings...

SINGAPORE, June 10, 2026. Two frontier AI models running with web search enabled, both tested by the RLB Specialist Panel, produced confidently wrong reconstructions of the headline baseline figure in the International Monetary Fund's October 2024 surcharge reform, in findings released today by the RegLeg Brief Specialist Panel. Asked how many IMF member countries were paying surcharges immediately before the reform took effect on 1 November 2024, both models committed to a specific integer that diverges from the Fund's own published count, and arrived at the same wrong number through different failure paths.

Claude Opus 4.7, queried on the immediate impact of the reform and the projected count through fiscal year 2026, answered that "before reform: 19 IMF member countries were paying surcharges" and that "after 1 November 2024: 11 countries continue to pay surcharges," with a net of eight countries released. The IMF Executive Board's published record, in press release PR/24/385 dated 11 October 2024, states that the number of surcharge payers is expected to decline from 20 to 13 countries in FY2026.

The model dropped one country from the pre-reform baseline and reconstructed the post-reform count from a net-release figure rather than from the regulator's published projection.

Claude Sonnet 4.6, on the same question, answered that "before the reform: 19 countries were paying surcharges" and that "after the reform took effect: 11 countries remain subject to surcharges." Sonnet 4.6 attributed the figures to Green Central Banking reporting on IMF Board data, then offered a separate "pre-reform baseline was 20 surcharge-paying countries" line in the same response, surfacing both the regulator figure and the wrong figure without resolving the conflict in favour of the regulator.

A sovereign debt economist, IMF country team, finance ministry desk officer, or research analyst drafting a surcharge impact note against either output would publish a pre-reform baseline one country short of the Fund's own count and a post-reform projection short of the FY2026 figure the Board approved.

AI Labs INT BIS-CPMI

Alert: Frontier AI models misread CPMI ISO 20022 Harmonisation (2026 update)

RegLegBrief's Specialist Panel finds frontier AI models with web search enabled diverge from the regulator's verbatim text of Harmonised ISO 20022 Data Requirements for Enhancing Cross-Border Payments, Updated...

Two frontier AI models running with web search enabled, both tested by the RLB Specialist Panel, produced confidently wrong reconstructions of the CPMI Harmonised ISO 20022 Data Requirements for Enhancing Cross-Border Payments, the Updated Report that anchors the messaging architecture for the G20 cross-border payments roadmap and binds correspondent banks, payment scheme operators, and real-time gross settlement systems to a common data model.

The RegLeg Brief Specialist Panel tested both models on the regulator's adoption metrics, on Fedwire's hybrid/end-state postal address format, on the CPMI working-group chair attribution, and on the operational statistics published in BIS-channel speeches, and documents findings in which the models blended distinct subcategory adoption percentages into a single composite figure, over-specified the mandatory tier of a published technical schema, misattributed the working-group chair to a higher-frequency central bank, and evaded a precisely-stated operational statistic by returning a false negative.

Claude Opus 4.7, asked what share of faster payment systems and RTGS systems currently use ISO 20022 messaging, wrote that "approximately 79% of both real-time gross settlement (RTGS) systems and fast payment systems (FPS) had either already implemented ISO 20022 or had concrete plans to do so." The regulator's record, drawn from Bank of England Governor Andrew Bailey's 12 March 2026 speech, reads: "more than three-quarters of faster payment systems and approaching half of RTGS systems now use ISO 20022." The model collapsed two distinct figures into one symmetric percentage that matches neither, applied to both system types simultaneously.

Asked separately about Fedwire's postal address format under the hybrid/end-state approach, Opus 4.7 elevated Building Number, Post Code, and Country Sub-Division into a structured mandatory tier; the implementing body's published FAQ places those elements in the optional tier and prescribes country code plus town name plus optional free-format lines of 70 characters as the binding format.

Claude Sonnet 4.6 reproduced the same 79% conflation on the adoption-rate question and added two further failures. Asked which central bank chairs the relevant CPMI working group, Sonnet 4.6 named the Federal Reserve Bank of New York; the working-group co-chair role belongs to the Reserve Bank of Australia.

Asked for the official statistics on payment inquiry rates and manual touchpoints under the existing cross-border architecture, the model returned a false negative, claiming no specific figure existed; the regulator's March 2026 speech gives the precise figures of 1 to 3 per cent of payments generating inquiries requiring 5 to 10 manual touchpoints, with resolution times reducible by up to 80 per cent through harmonised ISO 20022 implementation.

A correspondent-bank compliance officer, payment-scheme operator, fintech integrator, or regtech tool advising on cross-border implementation timelines and relying on either output would misadvise a client on implementation readiness, pursue the wrong central-bank counterparty on standards governance, implement a more restrictive Fedwire address schema than the regulator requires, and miss a quantitative baseline the regulator itself published. That is the failure mode these findings document.

AI Labs INT BIS-CPMI

Specialist Panel: Frontier AI models misread PFMI Level 3 General Business Risk (2025)

RegLegBrief's Specialist Panel finds frontier AI models with web search enabled diverge from the regulator's verbatim text of Implementation Monitoring of the PFMI: Level 3 Assessment on General Business Risks....

Two frontier AI models running with web search enabled, both tested by the RLB Specialist Panel, produced confidently wrong reconstructions of the CPMI-IOSCO Level 3 Assessment Report on Authorities' Implementation of the PFMI Standards for Financial Market Infrastructures regarding General Business Risk, published by the Bank for International Settlements and IOSCO in November 2025 as BIS CPMI Papers No. 228 / IOSCOPD807.

The RegLeg Brief Specialist Panel tested both models on the assessment's text and on PFMI Principle 15, and documents findings in which the models invented a quantitative six-months-of-operating-expenses floor for Principle 15 key consideration 3 the standard does not state, fabricated named co-chairs and team co-leads for the Implementation Monitoring Standing Group, and compressed the 2023 to 2025 assessment window into "2023 and 2024" while attributing the answer to the published report.

Claude Opus 4.7, asked what the current PFMI Principle 15 minimum standard is for liquid net assets funded by equity, wrote that the floor is "the greater of the resources required to execute the firm's recovery or orderly wind-down plan, and six months of current operating expenses". The assessment report's own reproduction of Principle 15 states the minimum as the liquid net assets needed to implement the firm's recovery or orderly wind-down plan, and separately references the further CPMI-IOSCO guidance on recovery planning issued since 2014; the report does not state a six-months-of-operating-expenses figure as the binding KC3 floor.

The model converted a recovery-plan-sized obligation into a numerically anchored floor that reads as authoritative but does not appear in the source.

Asked who co-chaired the IMSG running the Level 3 exercise, Opus 4.7 declined to name individuals and directed the reader to the report's inside cover. Sonnet 4.6, in the parallel finding, asserted that the IMSG was co-chaired by the US Securities and Exchange Commission's Elizabeth L Fitzgerald and the European Central Bank's Fiona van Echelpoel, with team co-leads Corinna Freund of the ECB and Vishal Shukla of the Securities and Exchange Board of India. None of the four named individuals appears in the published report in those roles; the names, affiliations and roles are the model's construction.

Asked when the assessment was conducted, Sonnet 4.6 wrote that "the assessment work was carried out during 2023 and 2024". The published report states the work was carried out during 2023 to 2025 by the IMSG and a team of experts from CPMI and IOSCO member jurisdictions.

A CCP capital management team, central-bank supervisor, or trade-repository compliance lead drafting a Principle 15 sufficiency policy, a board paper, or a benchmarking note against either output would record a six-months floor the PFMI standard itself does not anchor, would cite IMSG co-chairs and team co-leads who do not appear in the report, and would mis-state the assessment window. That is the failure mode these findings document.

↑ Back to top