AI Hallucination Research › Briefings

Briefings Blog

The running blog from the RLB Specialist Panel delves into real-world scenarios where the compliance, legal, or AI lab team interacts with frontier AI models under specific regulations. The blogs are anonymised to remove client-specific details and include insights from the RLB team analysing the hallucinations experienced in AI models while working on these cases. For example, when a model returns a confident answer that contradicts the regulator's primary text, such as a fabricated staff letter, a wrong appendix, or an inverted scope, these issues are discussed here. Each blog explains one set of findings and what it would have meant for the team that would have acted on it, sans this research initiative. This blog is frequently updated, a few times a day.

263 briefings in the archive · Subscribe via Atom: /briefings/feed.xml (this blog) · /feed.xml (all RegLegBrief publications)
Audience colours: AI Labs Practitioner (profession) Sector × Department
Audience
Jur.
Regulator
Profession
Sector
Dept
Range
Sort
Per page
Showing 5 of 263 · page 27 of 53
Monday, 06 July 2026
Practitioner: Lawyers US CFTC

Lawyers: AI summaries of CFTC Digital Asset Collateral & Tokenized Assets Staff Guidance (2025) may understate professional obligations

For Lawyers working with CFTC Digital Asset Collateral No-Action Relief and Tokenized Asset Staff Guidance (Market Participants Division, December 2025): where Specialist-Panel-verified divergences between frontier...

Lawyers advising on the CFTC Digital Asset Collateral Framework are increasingly using AI to draft 2-page client memos on payment stablecoin eligibility, generate partner-level briefings on the phased onboarding obligations for futures commission merchants, and validate staff-letter citation language against the published CFTC text before issuing legal opinions on customer margin collateral acceptance.

The RLB Specialist Panel put a set of practitioner-grade questions on the CFTC Digital Asset Collateral Framework to two frontier AI models with web search active. Each question is prepared by the Panel based on the workflows that lawyers actually use AI for under the Market Participants Division's December 2025 staff letter, as amended by Staff Letter 26-05. The Panel then binds every AI response to verbatim regulator-issued source text held as primary substrate.

On the CFTC Digital Asset Collateral Framework, the AI subjects returned three hallucinated answers for lawyers, in the form of Inverted-Position Fabrication, Dropped-Qualifier Misattribution, and Dropped-Qualifier Misstated Rule.

For lawyers issuing legal opinions, client memos, transactional documents, and regulatory submissions that engage the CFTC Digital Asset Collateral Framework, staff-letter citation accuracy is load-bearing: a counterparty, opposing counsel, or regulator who can identify a citation error or a missing cross-reference on first reading of the document calls the entire piece of advice into question.

An AI-drafted memo that classifies the weekly digital asset reporting obligation as sunsetting when the regulator continues it, or that describes payment stablecoin eligibility without the OCC Interpretive Letter 1183 hook, or that presents the base 20 per cent haircut as the multi-DCO rule, leaves the lawyer exposed to professional liability, the firm exposed to reputational risk, and the FCM or stablecoin issuer client exposed to a reporting violation, an eligibility defect, or a customer collateral shortfall.

The published Specialist Panel findings carry the following citation identifiers:

Sector: Mainboard / Premium-Listed Issuers and Dept: Legal GB FCA

Mainboard / Premium-Listed Issuers Legal teams: documentation and reporting gaps possible from AI reading of FCA Consumer Duty (PS22/9)

For Mainboard / Premium-Listed Issuers Legal teams working with Consumer Duty (PS22/9 + PRIN 2A): Specialist-Panel-verified findings on where AI summaries diverge from the regulator's text, and what that means for...

In-house legal teams at mainboard / premium-listed issuers carrying retail-eligible products under the Consumer Duty are increasingly using AI to draft scope opinions on the boundary between in-scope retail-customer products and out-of-scope group-insurance, large-risk, and reinsurance arrangements, and to validate FSMA-statute citations in board papers and Part 6 disclosure drafting. The work product feeds into prospectus drafting, listing-rule compliance memos, and director attestations.

Two frontier AI models tested by the RLB Specialist Panel produced 2 substantive failures on this regulation under audit conditions. The failure classes recorded are: Misstated Statutory Architecture, Reversed the PRIN 2A Group-Insurance Exclusion. Questions were prepared by the RLB Specialist Panel based on real practical AI usage in the workflows the respective audience uses AI for, and each finding is bound to verbatim regulator-issued source text held as primary substrate.

The Consumer Duty (PS22/9 introducing Principle 12 and PRIN 2A, in force for open products from 31 July 2023 and for closed products from 31 July 2024) is the central retail-conduct regime the FCA now uses to grade firm behaviour, and the failure modes seen here all land inside the day-to-day work product that mainboard-issuer legal teams sign off on.

For mainboard / premium-listed issuer legal teams, the operational consequence is direct. Prospectus drafting, listing-rule compliance memos, board legal opinions, and director attestations on Consumer Duty applicability all rest on accurate PRIN 2A scope framing and statutory-architecture citation. A defect imported from AI work product surfaces on listing-rule review or audit-committee challenge, and the in-house function carries the professional exposure.

Citation IDs for the findings in this brief: RLB-H-GB-FCA-CONSUMER-DUTY-PS22-9-Q002-Sonnet46, RLB-H-GB-FCA-CONSUMER-DUTY-PS22-9-Q018-Opus47. Each citation links to the per-finding record, the AI subject answer, and the regulator-issued substrate excerpt the answer was tested against. The RLB Specialist Panel maintains an audit-traceable record of which model produced which answer, against which substrate passage, and the binding is what makes the finding referenceable in firm work product and in supervisory correspondence.

The findings below are the ones that mainboard-issuer legal teams working under the Consumer Duty are most likely to encounter in the AI tools they already use, and the briefing sections that follow read each finding against the regulator-issued text.

Sector: Corporate Banking and Dept: Legal GB FCA

Corporate Banking Legal teams: documentation and reporting gaps possible from AI reading of FCA Consumer Duty (PS22/9)

For Corporate Banking Legal teams working with Consumer Duty (PS22/9 + PRIN 2A): Specialist-Panel-verified findings on where AI summaries diverge from the regulator's text, and what that means for the sector's...

In-house legal teams at corporate banks operating under the Consumer Duty are increasingly using AI to draft scope opinions on the boundary between in-scope SME retail customers and out-of-scope large-risk commercial contracts, validate group-insurance distribution scope, and prepare director briefings on PRIN 2A. The work product feeds into legal-opinion files for the SME segment and director attestations on Consumer Duty applicability.

Two frontier AI models tested by the RLB Specialist Panel produced 2 substantive failures on this regulation under audit conditions. The failure classes recorded are: Misstated Statutory Architecture, Reversed the PRIN 2A Group-Insurance Exclusion. Questions were prepared by the RLB Specialist Panel based on real practical AI usage in the workflows the respective audience uses AI for, and each finding is bound to verbatim regulator-issued source text held as primary substrate.

The Consumer Duty (PS22/9 introducing Principle 12 and PRIN 2A, in force for open products from 31 July 2023 and for closed products from 31 July 2024) is the central retail-conduct regime the FCA now uses to grade firm behaviour, and the failure modes seen here all land inside the day-to-day work product that corporate-banking in-house legal teams sign off on.

For corporate-banking legal, the operational consequence is direct. Scope opinions on the SME-versus-large-risk boundary, group-insurance scope memos, and director attestations on Consumer Duty applicability all rest on accurate PRIN 2A scope framing. A defect imported from AI work product surfaces on legal-file review or board challenge, and the in-house function carries the professional exposure.

Citation IDs for the findings in this brief: RLB-H-GB-FCA-CONSUMER-DUTY-PS22-9-Q002-Sonnet46, RLB-H-GB-FCA-CONSUMER-DUTY-PS22-9-Q018-Opus47. Each citation links to the per-finding record, the AI subject answer, and the regulator-issued substrate excerpt the answer was tested against. The RLB Specialist Panel maintains an audit-traceable record of which model produced which answer, against which substrate passage, and the binding is what makes the finding referenceable in firm work product and in supervisory correspondence.

The findings below are the ones that corporate-banking in-house legal teams working under the Consumer Duty are most likely to encounter in the AI tools they already use, and the briefing sections that follow read each finding against the regulator-issued text.

Sector: Corporate Banking and Dept: Compliance GB FCA

Corporate Banking Compliance teams: documentation and reporting gaps possible from AI reading of FCA Consumer Duty (PS22/9)

For Corporate Banking Compliance teams working with Consumer Duty (PS22/9 + PRIN 2A): Specialist-Panel-verified findings on where AI summaries diverge from the regulator's text, and what that means for the sector's...

Compliance officers at corporate banks operating under the Consumer Duty are increasingly using AI to validate the scope boundary between in-scope SME retail customers and out-of-scope large-risk commercial customers, update Dear CEO letter expectation registers in light of FS25/2, and reconcile FCA Feedback Statements against existing supervisory correspondence on relationship-managed accounts. The work product feeds into the bank's compliance monitoring plan and the SME-segment supervisory dialogue.

Two frontier AI models tested by the RLB Specialist Panel produced 3 substantive failures on this regulation under audit conditions. The failure classes recorded are: Hedge in Place of Verified FS25/2 Figure, Refusal to Confirm a Documented FS25/2 Count, Invented Dual-Event Timeline for a Single FS25/2 Withdrawal. Questions were prepared by the RLB Specialist Panel based on real practical AI usage in the workflows the respective audience uses AI for, and each finding is bound to verbatim regulator-issued source text held as primary substrate.

The Consumer Duty (PS22/9 introducing Principle 12 and PRIN 2A, in force for open products from 31 July 2023 and for closed products from 31 July 2024) is the central retail-conduct regime the FCA now uses to grade firm behaviour, and the failure modes seen here all land inside the day-to-day work product that corporate-banking compliance teams sign off on.

For corporate-banking compliance, the operational consequence is direct. The compliance monitoring plan, the SME-segment supervisory dialogue, and the Dear CEO letter expectation register all rest on accurate framing of scope boundaries and of recent FCA Feedback Statements such as FS25/2. A defect imported from AI work product surfaces on supervisory follow-up, and the function carries the regulatory exposure.

Citation IDs for the findings in this brief: RLB-H-GB-FCA-CONSUMER-DUTY-PS22-9-Q013-Opus47, RLB-H-GB-FCA-CONSUMER-DUTY-PS22-9-Q013-Sonnet46, RLB-H-GB-FCA-CONSUMER-DUTY-PS22-9-Q020-Opus47. Each citation links to the per-finding record, the AI subject answer, and the regulator-issued substrate excerpt the answer was tested against. The RLB Specialist Panel maintains an audit-traceable record of which model produced which answer, against which substrate passage, and the binding is what makes the finding referenceable in firm work product and in supervisory correspondence.

The findings below are the ones that corporate-banking compliance teams working under the Consumer Duty are most likely to encounter in the AI tools they already use, and the briefing sections that follow read each finding against the regulator-issued text.

Sunday, 05 July 2026
Sector: Payment Institutions and Dept: Risk GB FCA

Payment Institutions Risk teams: documentation and reporting gaps possible from AI reading of FCA Consumer Duty (PS22/9)

For Payment Institutions Risk teams working with Consumer Duty (PS22/9 + PRIN 2A): Specialist-Panel-verified findings on where AI summaries diverge from the regulator's text, and what that means for the sector's...

Risk teams at payment institutions and e-money firms operating under the Consumer Duty are increasingly using AI to update foreseeable-harm risk matrices for retail-customer journeys, validate fair-value risk assessments for new products, and stress-test the firm's customer-outcome KPIs against PRIN 2A. The work product feeds directly into the firm's risk register and the executive-risk-committee dashboard.

Two frontier AI models tested by the RLB Specialist Panel produced 3 substantive failures on this regulation under audit conditions. The failure classes recorded are: Inference Drift on the Foreseeable-Harm Safe Harbour, Inference Drift on Fair Value Quantification Expectation, Inference Drift on Required Depth of Non-Monetary Analysis. Questions were prepared by the RLB Specialist Panel based on real practical AI usage in the workflows the respective audience uses AI for, and each finding is bound to verbatim regulator-issued source text held as primary substrate.

The Consumer Duty (PS22/9 introducing Principle 12 and PRIN 2A, in force for open products from 31 July 2023 and for closed products from 31 July 2024) is the central retail-conduct regime the FCA now uses to grade firm behaviour, and the failure modes seen here all land inside the day-to-day work product that payment-institutions risk teams sign off on.

For payment-institutions risk, the operational consequence is direct. The risk register, the foreseeable-harm matrix for retail-customer journeys, and the executive-risk-committee dashboard all rest on accurate PRIN 2A framing. A defect imported from AI work product surfaces on internal-audit pull or supervisor review, and the risk function carries the second-line exposure.

Citation IDs for the findings in this brief: RLB-H-GB-FCA-CONSUMER-DUTY-PS22-9-Q003-Opus47, RLB-H-GB-FCA-CONSUMER-DUTY-PS22-9-Q008-Opus47, RLB-H-GB-FCA-CONSUMER-DUTY-PS22-9-Q008-Sonnet46. Each citation links to the per-finding record, the AI subject answer, and the regulator-issued substrate excerpt the answer was tested against. The RLB Specialist Panel maintains an audit-traceable record of which model produced which answer, against which substrate passage, and the binding is what makes the finding referenceable in firm work product and in supervisory correspondence.

The findings below are the ones that payment-institutions risk teams working under the Consumer Duty are most likely to encounter in the AI tools they already use, and the briefing sections that follow read each finding against the regulator-issued text.

↑ Back to top