AI Hallucination Research › Briefings

Briefings Blog

The running blog from the RLB Specialist Panel delves into real-world scenarios where the compliance, legal, or AI lab team interacts with frontier AI models under specific regulations. The blogs are anonymised to remove client-specific details and include insights from the RLB team analysing the hallucinations experienced in AI models while working on these cases. For example, when a model returns a confident answer that contradicts the regulator's primary text, such as a fabricated staff letter, a wrong appendix, or an inverted scope, these issues are discussed here. Each blog explains one set of findings and what it would have meant for the team that would have acted on it, sans this research initiative. This blog is frequently updated, a few times a day.

263 briefings in the archive · Subscribe via Atom: /briefings/feed.xml (this blog) · /feed.xml (all RegLegBrief publications)
Audience colours: AI Labs Practitioner (profession) Sector × Department
Audience
Jur.
Regulator
Profession
Sector
Dept
Range
Sort
Per page
Showing 5 of 263 · page 28 of 53
Sunday, 05 July 2026
Sector: Payment Institutions and Dept: Risk GB FCA

Payment Institutions Risk teams: documentation and reporting gaps possible from AI reading of FCA Consumer Duty (PS22/9)

For Payment Institutions Risk teams working with Consumer Duty (PS22/9 + PRIN 2A): Specialist-Panel-verified findings on where AI summaries diverge from the regulator's text, and what that means for the sector's...

Risk teams at payment institutions and e-money firms operating under the Consumer Duty are increasingly using AI to update foreseeable-harm risk matrices for retail-customer journeys, validate fair-value risk assessments for new products, and stress-test the firm's customer-outcome KPIs against PRIN 2A. The work product feeds directly into the firm's risk register and the executive-risk-committee dashboard.

Two frontier AI models tested by the RLB Specialist Panel produced 3 substantive failures on this regulation under audit conditions. The failure classes recorded are: Inference Drift on the Foreseeable-Harm Safe Harbour, Inference Drift on Fair Value Quantification Expectation, Inference Drift on Required Depth of Non-Monetary Analysis. Questions were prepared by the RLB Specialist Panel based on real practical AI usage in the workflows the respective audience uses AI for, and each finding is bound to verbatim regulator-issued source text held as primary substrate.

The Consumer Duty (PS22/9 introducing Principle 12 and PRIN 2A, in force for open products from 31 July 2023 and for closed products from 31 July 2024) is the central retail-conduct regime the FCA now uses to grade firm behaviour, and the failure modes seen here all land inside the day-to-day work product that payment-institutions risk teams sign off on.

For payment-institutions risk, the operational consequence is direct. The risk register, the foreseeable-harm matrix for retail-customer journeys, and the executive-risk-committee dashboard all rest on accurate PRIN 2A framing. A defect imported from AI work product surfaces on internal-audit pull or supervisor review, and the risk function carries the second-line exposure.

Citation IDs for the findings in this brief: RLB-H-GB-FCA-CONSUMER-DUTY-PS22-9-Q003-Opus47, RLB-H-GB-FCA-CONSUMER-DUTY-PS22-9-Q008-Opus47, RLB-H-GB-FCA-CONSUMER-DUTY-PS22-9-Q008-Sonnet46. Each citation links to the per-finding record, the AI subject answer, and the regulator-issued substrate excerpt the answer was tested against. The RLB Specialist Panel maintains an audit-traceable record of which model produced which answer, against which substrate passage, and the binding is what makes the finding referenceable in firm work product and in supervisory correspondence.

The findings below are the ones that payment-institutions risk teams working under the Consumer Duty are most likely to encounter in the AI tools they already use, and the briefing sections that follow read each finding against the regulator-issued text.

Sector: Payment Institutions and Dept: Legal GB FCA

Payment Institutions Legal teams: documentation and reporting gaps possible from AI reading of FCA Consumer Duty (PS22/9)

For Payment Institutions Legal teams working with Consumer Duty (PS22/9 + PRIN 2A): Specialist-Panel-verified findings on where AI summaries diverge from the regulator's text, and what that means for the sector's...

In-house legal teams at payment institutions and e-money firms operating under the Consumer Duty are increasingly using AI to validate Principle 12 scope opinions for retail-customer-facing services, draft scope memos on PRIN 2A exclusions, and prepare director briefings on FCA Feedback Statements such as FS25/2. The work product sits at the centre of new-product legal opinions, scope-of-application memos, and supervisor-correspondence drafting.

Two frontier AI models tested by the RLB Specialist Panel produced 4 substantive failures on this regulation under audit conditions. The failure classes recorded are: Misstated Statutory Architecture, Reversed the PRIN 2A Group-Insurance Exclusion, Invented Dual-Event Timeline for a Single FS25/2 Withdrawal, Refusal to Confirm FS25/2 Withdrawal Count. Questions were prepared by the RLB Specialist Panel based on real practical AI usage in the workflows the respective audience uses AI for, and each finding is bound to verbatim regulator-issued source text held as primary substrate.

The Consumer Duty (PS22/9 introducing Principle 12 and PRIN 2A, in force for open products from 31 July 2023 and for closed products from 31 July 2024) is the central retail-conduct regime the FCA now uses to grade firm behaviour, and the failure modes seen here all land inside the day-to-day work product that payment-institutions in-house legal teams sign off on.

For payment-institutions legal, the operational consequence is direct. Scope-of-application memos, new-product legal opinions, and director attestations on Consumer Duty applicability all rest on accurate PRIN 2A scope and FSMA-statute framing. A defect imported from AI work product surfaces on legal-file review or board challenge, and the in-house function carries the professional exposure.

Citation IDs for the findings in this brief: RLB-H-GB-FCA-CONSUMER-DUTY-PS22-9-Q002-Sonnet46, RLB-H-GB-FCA-CONSUMER-DUTY-PS22-9-Q018-Opus47, RLB-H-GB-FCA-CONSUMER-DUTY-PS22-9-Q020-Opus47, RLB-H-GB-FCA-CONSUMER-DUTY-PS22-9-Q020-Sonnet46. Each citation links to the per-finding record, the AI subject answer, and the regulator-issued substrate excerpt the answer was tested against. The RLB Specialist Panel maintains an audit-traceable record of which model produced which answer, against which substrate passage, and the binding is what makes the finding referenceable in firm work product and in supervisory correspondence.

The findings below are the ones that payment-institutions in-house legal teams working under the Consumer Duty are most likely to encounter in the AI tools they already use, and the briefing sections that follow read each finding against the regulator-issued text.

Sector: Retail Banking and Dept: Risk GB FCA

Retail Banking Risk teams: documentation and reporting gaps possible from AI reading of FCA Consumer Duty (PS22/9)

For Retail Banking Risk teams working with Consumer Duty (PS22/9 + PRIN 2A): Specialist-Panel-verified findings on where AI summaries diverge from the regulator's text, and what that means for the sector's...

Risk teams at retail banks operating under the Consumer Duty are increasingly using AI to update foreseeable-harm matrices, draft fair-value risk assessments for product committees, and validate the bank's customer-outcome KPIs against the FCA's stated expectations. The work product feeds directly into the bank's Consumer Duty risk register and the executive-risk-committee dashboard the supervisor sees.

Two frontier AI models tested by the RLB Specialist Panel produced 4 substantive failures on this regulation under audit conditions. The failure classes recorded are: Inference Drift on the Foreseeable-Harm Safe Harbour, Confused Guidance with Rule on Consumer Testing, Inference Drift on Fair Value Quantification Expectation, Inference Drift on Required Depth of Non-Monetary Analysis. Questions were prepared by the RLB Specialist Panel based on real practical AI usage in the workflows the respective audience uses AI for, and each finding is bound to verbatim regulator-issued source text held as primary substrate.

The Consumer Duty (PS22/9 introducing Principle 12 and PRIN 2A, in force for open products from 31 July 2023 and for closed products from 31 July 2024) is the central retail-conduct regime the FCA now uses to grade firm behaviour, and the failure modes seen here all land inside the day-to-day work product that retail-banking risk teams sign off on.

For retail-banking risk, the operational consequence is direct. The Consumer Duty risk register, the foreseeable-harm matrix, and the executive-risk-committee dashboard all rest on accurate PRIN 2A framing. A defect imported from AI work product surfaces in supervisor dashboard review or internal-audit, and the risk function carries the second-line exposure.

Citation IDs for the findings in this brief: RLB-H-GB-FCA-CONSUMER-DUTY-PS22-9-Q003-Opus47, RLB-H-GB-FCA-CONSUMER-DUTY-PS22-9-Q007-Sonnet46, RLB-H-GB-FCA-CONSUMER-DUTY-PS22-9-Q008-Opus47, RLB-H-GB-FCA-CONSUMER-DUTY-PS22-9-Q008-Sonnet46. Each citation links to the per-finding record, the AI subject answer, and the regulator-issued substrate excerpt the answer was tested against. The RLB Specialist Panel maintains an audit-traceable record of which model produced which answer, against which substrate passage, and the binding is what makes the finding referenceable in firm work product and in supervisory correspondence.

The findings below are the ones that retail-banking risk teams working under the Consumer Duty are most likely to encounter in the AI tools they already use, and the briefing sections that follow read each finding against the regulator-issued text.

Sector: Payment Institutions and Dept: Compliance GB FCA

Payment Institutions Compliance teams: documentation and reporting gaps possible from AI reading of FCA Consumer Duty (PS22/9)

For Payment Institutions Compliance teams working with Consumer Duty (PS22/9 + PRIN 2A): Specialist-Panel-verified findings on where AI summaries diverge from the regulator's text, and what that means for the...

Compliance officers at payment institutions and e-money firms operating under the Consumer Duty are increasingly using AI to validate retail-customer scope analyses, update foreseeable-harm monitoring against transaction patterns, draft summaries of FCA Feedback Statements such as FS25/2, and reconcile Dear CEO letter retirements against existing supervisory expectation registers. The work product feeds directly into the firm's compliance monitoring plan and the supervisor's annual relationship correspondence.

Two frontier AI models tested by the RLB Specialist Panel produced 5 substantive failures on this regulation under audit conditions. The failure classes recorded are: Inference Drift on the Foreseeable-Harm Safe Harbour, Confused Guidance with Rule on Consumer Testing, Hedge in Place of Verified FS25/2 Figure, Invented Dual-Event Timeline for a Single FS25/2 Withdrawal, Refusal to Confirm FS25/2 Withdrawal Count. Questions were prepared by the RLB Specialist Panel based on real practical AI usage in the workflows the respective audience uses AI for, and each finding is bound to verbatim regulator-issued source text held as primary substrate.

The Consumer Duty (PS22/9 introducing Principle 12 and PRIN 2A, in force for open products from 31 July 2023 and for closed products from 31 July 2024) is the central retail-conduct regime the FCA now uses to grade firm behaviour, and the failure modes seen here all land inside the day-to-day work product that payment-institutions compliance teams sign off on.

For payment-institutions compliance, the operational consequence is direct. The compliance monitoring plan, the annual board report on Consumer Duty, and the supervisor's annual relationship correspondence all rest on accurate framing of the rule and of recent FCA Feedback Statements. A defect imported from AI work product surfaces on supervisory follow-up, and the function carries the regulatory exposure.

Citation IDs for the findings in this brief: RLB-H-GB-FCA-CONSUMER-DUTY-PS22-9-Q003-Opus47, RLB-H-GB-FCA-CONSUMER-DUTY-PS22-9-Q007-Sonnet46, RLB-H-GB-FCA-CONSUMER-DUTY-PS22-9-Q013-Opus47, RLB-H-GB-FCA-CONSUMER-DUTY-PS22-9-Q020-Opus47, RLB-H-GB-FCA-CONSUMER-DUTY-PS22-9-Q020-Sonnet46. Each citation links to the per-finding record, the AI subject answer, and the regulator-issued substrate excerpt the answer was tested against. The RLB Specialist Panel maintains an audit-traceable record of which model produced which answer, against which substrate passage, and the binding is what makes the finding referenceable in firm work product and in supervisory correspondence.

The findings below are the ones that payment-institutions compliance teams working under the Consumer Duty are most likely to encounter in the AI tools they already use, and the briefing sections that follow read each finding against the regulator-issued text.

Sector: Retail Banking and Dept: Product & Business Development GB FCA

Retail Banking Product & Business Development teams: documentation and reporting gaps possible from AI reading of FCA Consumer Duty (PS22/9)

For Retail Banking Product & Business Development teams working with Consumer Duty (PS22/9 + PRIN 2A): Specialist-Panel-verified findings on where AI summaries diverge from the regulator's text, and what that means...

Product and business development teams at retail banks operating under the Consumer Duty are increasingly using AI to validate fair-value rationales for new-product approvals, draft target-market-statement language for product-governance files, and prepare go-to-market briefings that map customer-outcome design choices to PRIN 2A.4. The work product feeds directly into the bank's product-governance approval packs and the post-launch monitoring evidence the supervisor reviews.

Two frontier AI models tested by the RLB Specialist Panel produced 5 substantive failures on this regulation under audit conditions. The failure classes recorded are: Inference Drift on the Foreseeable-Harm Safe Harbour, Confused Guidance with Rule on Consumer Testing, Inference Drift on Fair Value Quantification Expectation, Inference Drift on Required Depth of Non-Monetary Analysis, Reversed the PRIN 2A Group-Insurance Exclusion. Questions were prepared by the RLB Specialist Panel based on real practical AI usage in the workflows the respective audience uses AI for, and each finding is bound to verbatim regulator-issued source text held as primary substrate.

The Consumer Duty (PS22/9 introducing Principle 12 and PRIN 2A, in force for open products from 31 July 2023 and for closed products from 31 July 2024) is the central retail-conduct regime the FCA now uses to grade firm behaviour, and the failure modes seen here all land inside the day-to-day work product that retail-banking product and business-development teams sign off on.

For retail-banking product and business-development teams, the operational consequence is direct. Product-governance approval packs, target-market statements, and post-launch monitoring evidence all rest on accurate fair-value and PRIN 2A.4 framing. A defect imported from AI work product surfaces on product-board re-review or thematic supervision, and the product function carries the launch-risk exposure.

Citation IDs for the findings in this brief: RLB-H-GB-FCA-CONSUMER-DUTY-PS22-9-Q003-Opus47, RLB-H-GB-FCA-CONSUMER-DUTY-PS22-9-Q007-Sonnet46, RLB-H-GB-FCA-CONSUMER-DUTY-PS22-9-Q008-Opus47, RLB-H-GB-FCA-CONSUMER-DUTY-PS22-9-Q008-Sonnet46, RLB-H-GB-FCA-CONSUMER-DUTY-PS22-9-Q018-Opus47. Each citation links to the per-finding record, the AI subject answer, and the regulator-issued substrate excerpt the answer was tested against. The RLB Specialist Panel maintains an audit-traceable record of which model produced which answer, against which substrate passage, and the binding is what makes the finding referenceable in firm work product and in supervisory correspondence.

The findings below are the ones that retail-banking product and business-development teams working under the Consumer Duty are most likely to encounter in the AI tools they already use, and the briefing sections that follow read each finding against the regulator-issued text.

↑ Back to top