AI Hallucination Research › Briefings

Briefings Blog

The running blog from the RLB Specialist Panel delves into real-world scenarios where the compliance, legal, or AI lab team interacts with frontier AI models under specific regulations. The blogs are anonymised to remove client-specific details and include insights from the RLB team analysing the hallucinations experienced in AI models while working on these cases. For example, when a model returns a confident answer that contradicts the regulator's primary text, such as a fabricated staff letter, a wrong appendix, or an inverted scope, these issues are discussed here. Each blog explains one set of findings and what it would have meant for the team that would have acted on it, sans this research initiative. This blog is frequently updated, a few times a day.

263 briefings in the archive · Subscribe via Atom: /briefings/feed.xml (this blog) · /feed.xml (all RegLegBrief publications)
Audience colours: AI Labs Practitioner (profession) Sector × Department
Audience
Jur.
Regulator
Profession
Sector
Dept
Range
Sort
Per page
Showing 5 of 263 · page 9 of 53
Tuesday, 21 July 2026
Practitioner: Accountants (CA/PA) INT BIS-CPMI

Accountants (CA/PA): AI summaries of CPMI-IOSCO VM Effective Practices 2025 may understate professional obligations

For Accountants (CA/PA) working with Streamlining Variation Margin in Centrally Cleared Markets — Examples of Effective Practices: where Specialist-Panel-verified divergences between frontier AI summaries and the...

Accountants supporting central counterparties, clearing members, and asset-management clients on variation margin control reviews are increasingly using AI to map d226 effective practices against client control libraries, draft compliance scoring matrices for internal audit, prepare regulator-readiness self-assessments for CCP governance committees, and validate disclosure footnotes that reference CPMI-IOSCO standards. Leading AI assistants tested by the RLB Specialist Panel produced confident, citable answers on the binding force of d226 that the document itself directly contradicts.

The RLB Specialist Panel tested whether two frontier AI models could correctly characterise the legal status of d226, asking them to classify each of the eight effective practices set out in the document as either a mandatory requirement with enforcement consequences, a supervisory expectation that regulators will test against, or voluntary guidance with no binding legal force. The exercise targeted what the Panel calls inverted modality: AI commitments that flip the binding force of a source text from voluntary illustration to supervisory or mandatory rule.

The frontier model under test produced a complete compliance obligations memo that classified every one of the eight effective practices as either a supervisory expectation in its own right or as overlapping with mandatory national rules, with a threshold classification asserting that d226 carries "a strong gravitational pull into (B) SUPERVISORY EXPECTATION." The document's own stated purpose paragraph, by contrast, records that d226 sets out "examples of how standards set out in the CPMI-IOSCO Principles for financial market infrastructures, as supplemented by the relevant guidance, can be met."

For accountants, the operational consequence is direct. Control assessment matrices and audit work papers that score clients against d226 practices as if they were supervisory pass/fail criteria distort residual-risk ratings, push remediation budgets toward voluntary illustrations, and risk under-rating exposures that arise from the underlying PFMI Principles or national rulebook overlays. The pattern is also reproducible: it surfaces wherever a deliverable asks the model to commit to a legal characterisation of an international standard-setter publication, and it is not addressed by general-purpose prompting.

The RLB Specialist Panel records the finding under the misstated-rule failure category and binds it to verbatim regulator text drawn from the d226 final report held as primary substrate.

The full finding is recorded under Citation ID RLB-H-INT-BIS-CPMI-CPMI-IOSCO-VARIATION-MARGIN-CCPs-2025-Q004-Opus47. The regulation hub is at /regulators/j1/INT/BIS-CPMI-INT-001/CPMI-IOSCO-VARIATION-MARGIN-CCPs-2025/. Questions are prepared by the RLB Specialist Panel based on real practical AI usage in the workflows the respective audience uses AI for. The Panel binds each AI finding to verbatim regulator-issued source text held as primary substrate.

Practitioner: Lawyers INT BIS-CPMI

Lawyers: AI summaries of CPMI-IOSCO VM Effective Practices 2025 may understate professional obligations

For Lawyers working with Streamlining Variation Margin in Centrally Cleared Markets — Examples of Effective Practices: where Specialist-Panel-verified divergences between frontier AI summaries and the regulator's...

Lawyers advising central counterparties, clearing members, and asset managers on variation margin obligations are increasingly using AI to draft board memos on CPMI-IOSCO d226, classify the legal status of each effective practice for the Audit and Risk Committee, validate cross-references between d226 and the underlying PFMI Principles, and prepare partner-level briefings on whether national regulators will treat the document as a supervisory expectation or as non-binding guidance. Leading AI assistants tested by the RLB Specialist Panel produced confident, citable answers on the binding force of d226 that the document itself directly contradicts.

The RLB Specialist Panel tested whether two frontier AI models could correctly characterise the legal status of d226, asking them to classify each of the eight effective practices set out in the document as either a mandatory requirement with enforcement consequences, a supervisory expectation that regulators will test against, or voluntary guidance with no binding legal force. The exercise targeted what the Panel calls inverted modality: AI commitments that flip the binding force of a source text from voluntary illustration to supervisory or mandatory rule.

The frontier model under test produced a complete compliance obligations memo that classified every one of the eight effective practices as either a supervisory expectation in its own right or as overlapping with mandatory national rules, with a threshold classification asserting that d226 carries "a strong gravitational pull into (B) SUPERVISORY EXPECTATION." The document's own stated purpose paragraph, by contrast, records that d226 sets out "examples of how standards set out in the CPMI-IOSCO Principles for financial market infrastructures, as supplemented by the relevant guidance, can be met."

For lawyers, the operational consequence is direct. Partner-level memos and board briefings that classify d226 practices as supervisory expectations or mandatory obligations create disclosure exposure to clients, audit committees, and counterparties who rely on the legal characterisation to size implementation budgets, draft rulebook amendments, and respond to supervisory dialogue. The pattern is also reproducible: it surfaces wherever a deliverable asks the model to commit to a legal characterisation of an international standard-setter publication, and it is not addressed by general-purpose prompting.

The RLB Specialist Panel records the finding under the misstated-rule failure category and binds it to verbatim regulator text drawn from the d226 final report held as primary substrate.

The full finding is recorded under Citation ID RLB-H-INT-BIS-CPMI-CPMI-IOSCO-VARIATION-MARGIN-CCPs-2025-Q004-Opus47. The regulation hub is at /regulators/j1/INT/BIS-CPMI-INT-001/CPMI-IOSCO-VARIATION-MARGIN-CCPs-2025/. Questions are prepared by the RLB Specialist Panel based on real practical AI usage in the workflows the respective audience uses AI for. The Panel binds each AI finding to verbatim regulator-issued source text held as primary substrate.

AI Labs INT BIS-CPMI

Alert: Frontier AI models misread CPMI-IOSCO VM Effective Practices 2025

RegLegBrief's Specialist Panel finds frontier AI models with web search enabled diverge from the regulator's verbatim text of Streamlining Variation Margin in Centrally Cleared Markets, Examples of Effective...

AI lab teams fielding frontier models into capital markets workflows are routinely asked to characterise the legal status of international standard-setter publications. The Bank for International Settlements' Committee on Payments and Market Infrastructures and the International Organization of Securities Commissions issued document d226 on 15 January 2025, setting out eight effective practices for streamlining variation margin in centrally cleared markets.

The document expressly records its own stated purpose as providing "examples of how standards set out in the CPMI-IOSCO Principles for financial market infrastructures, as supplemented by the relevant guidance, can be met." A frontier AI model tested by the RLB Specialist Panel returned a confident, citable compliance obligations memo that converted that voluntary illustration into a supervisory baseline.

The Specialist Panel ran a deliverable-pressure probe that placed the model in the role of a CCP General Counsel preparing a compliance obligations memo for the board's Audit and Risk Committee, with the deliverable required to classify each of the eight d226 effective practices as (A) mandatory requirement, (B) supervisory expectation, or (C) voluntary guidance. The probe surfaces what the Panel calls inverted modality: an AI commitment that flips the binding force of a source text from voluntary illustration to supervisory or mandatory rule under deliverable pressure.

The model produced a complete memo whose threshold paragraph correctly identified d226 as voluntary, then immediately overrode that identification, and proceeded to classify every one of the eight practices as a supervisory expectation in its own right or as overlapping with mandatory national rules.

For AI lab teams, the operational implication is that a frontier model under deliverable pressure on an international standard-setter publication may default to the more demanding characterisation even where the source text and the model's own threshold paragraph state otherwise. The failure pattern is reproducible, surfaces only under deliverable prompts that require specific per-item classifications and cited language, and is not addressed by general-purpose prompting. The RLB Specialist Panel records the finding under the misstated-rule failure category and binds it to verbatim regulator text drawn from the d226 final report held as primary substrate.

The full finding is recorded under Citation ID RLB-H-INT-BIS-CPMI-CPMI-IOSCO-VARIATION-MARGIN-CCPs-2025-Q004-Opus47. The regulation hub is at /regulators/j1/INT/BIS-CPMI-INT-001/CPMI-IOSCO-VARIATION-MARGIN-CCPs-2025/. Questions are prepared by the RLB Specialist Panel based on real practical AI usage in the workflows the respective audience uses AI for. The Panel binds each AI finding to verbatim regulator-issued source text held as primary substrate.

Monday, 20 July 2026
Sector: Telecommunications and Dept: Legal INT OECD

Telecommunications Legal teams: documentation and reporting gaps possible from AI reading of Recommendation of the Council on Merger Review

For Telecommunications Legal teams working with Recommendation of the Council on Merger Review (2025 Revision): Specialist-Panel-verified findings on where AI summaries diverge from the regulator's text, and what...

Legal teams at telecommunications groups approaching cross-border consolidation transactions under the 2025 OECD Merger Review Recommendation are increasingly using AI to draft regulatory-strategy memos on remedies hierarchy and structural-divestiture sequencing, generate executive-committee briefings on cross-border clearance exposure, and validate Section IV.3 remedies-priority language against the OECD text before remedy negotiations open with authorities.

The RLB Specialist Panel put a set of practitioner-grade questions on the 2025 OECD Merger Review Recommendation to two frontier AI models with web search active. Each question is prepared by the Panel based on the workflows that legal teams at telecommunications firms actually use AI for under the OECD's 2025 revision of the Recommendation of the Council on Merger Review (OECD/LEGAL/0333). The Panel then binds every AI response to verbatim regulator-issued source text held as primary substrate.

On the 2025 OECD Merger Review Recommendation, the AI subjects returned a single hallucinated answer for legal teams at telecommunications firms, in the form of Misattributed Cross-Jurisdictional Doctrine.

For legal teams at telecommunications firms advising on cross-border merger transactions touching the 2025 OECD Merger Review Recommendation, citation accuracy on the operative architecture, on Section IV.3 remedies hierarchy, and on Section III.11.b failing firm defence is load-bearing in every authority-facing submission, every board memo, and every transactional document. A counterparty or competition authority who identifies a structural inflation, a misattributed sub-hierarchy, or a closed-cumulative-test framing on first reading calls the entire piece of advice into question.

The structural-architecture failure is the most directly visible: a board memo or regulator-facing submission that lists 'international co-operation' or 'monitoring' as operative RECOMMENDS sections is wrong on first reading. The Section IV.3 EU sub-hierarchy import is the most insidious failure, reading as authoritative because the EU framework is real, but presenting EU practice as OECD content imports the wrong normative baseline into the firm's remedy strategy.

The published Specialist Panel findings carry the following citation identifiers:

Sector: Statutory Boards & Agencies and Dept: Compliance INT OECD

Statutory Boards & Agencies Compliance teams: documentation and reporting gaps possible from AI reading of Recommendation of the Council on Merger Review

For Statutory Boards & Agencies Compliance teams working with Recommendation of the Council on Merger Review (2025 Revision): Specialist-Panel-verified findings on where AI summaries diverge from the regulator's...

Compliance teams at statutory boards and agencies coordinating with competition authorities on cross-border merger-review reporting cycles and on the 2025 OECD Merger Review Recommendation are increasingly using AI to draft inter-agency memos on the Council-reporting cadence, generate engagement briefings on the Competition Committee monitoring cycle, prepare summaries of the Section V ex-post-assessment obligation for senior officials, and validate Section VIII.c reporting-interval language against the OECD text before inter-agency reporting cycles open.

The RLB Specialist Panel put a set of practitioner-grade questions on the 2025 OECD Merger Review Recommendation to two frontier AI models with web search active. Each question is prepared by the Panel based on the workflows that compliance teams at statutory boards & agencies actually use AI for under the OECD's 2025 revision of the Recommendation of the Council on Merger Review (OECD/LEGAL/0333). The Panel then binds every AI response to verbatim regulator-issued source text held as primary substrate.

On the 2025 OECD Merger Review Recommendation, the AI subjects returned a single hallucinated answer for compliance teams at statutory boards & agencies, in the form of Open-Interval Collapse.

For compliance teams at statutory boards & agencies coordinating inter-agency reporting cycles that engage the 2025 OECD Merger Review Recommendation, the Section VIII.c Council-reporting cadence and the Section V ex-post-assessment obligation drive the reporting-tracker design and the engagement-script for Competition Committee monitoring. A reporting tracker built on a fixed-cycle five-year cadence mis-schedules the second report, locks in a 2035 date the Recommendation does not set, and signals to the authority-side reviewer that the underlying regulatory map is unreliable.

The Section V ex-post-assessment obligation, the headline addition of the 2025 revision, is the obligation a compliance team would most want surfaced in its tracker; omitting it from the operative architecture is a substantive gap that the next inter-agency engagement will expose.

The published Specialist Panel findings carry the following citation identifiers:

↑ Back to top