AI Hallucination Research › Briefings

Briefings Blog

The running blog from the RLB Specialist Panel delves into real-world scenarios where the compliance, legal, or AI lab team interacts with frontier AI models under specific regulations. The blogs are anonymised to remove client-specific details and include insights from the RLB team analysing the hallucinations experienced in AI models while working on these cases. For example, when a model returns a confident answer that contradicts the regulator's primary text, such as a fabricated staff letter, a wrong appendix, or an inverted scope, these issues are discussed here. Each blog explains one set of findings and what it would have meant for the team that would have acted on it, sans this research initiative. This blog is frequently updated, a few times a day.

263 briefings in the archive · Subscribe via Atom: /briefings/feed.xml (this blog) · /feed.xml (all RegLegBrief publications)
Audience colours: AI Labs Practitioner (profession) Sector × Department
Audience
Jur.
Regulator
Profession
Sector
Dept
Range
Sort
Per page
Showing 5 of 263 · page 52 of 53
Sunday, 14 June 2026
AI Labs INT OECD

Specialist Panel: Frontier AI models misread Recommendation of the Council on Digital Technologies and the Environment

RegLegBrief's Specialist Panel finds frontier AI models with web search enabled diverge from the regulator's verbatim text of Recommendation of the Council on Digital Technologies and the Environment (2025 Revision)....

One frontier AI model with web search enabled, produced a confidently wrong figure for Ireland's 2021 data-centre share of national metered electricity, fabricating a 14% share where the regulator-cited verbatim text from Ireland's Central Statistics Office, as embedded in the OECD Digital Economy Outlook 2024 chapter, sets the figure at 11%. The Specialist Panel tested the model with an application-style probe bounded to substrate accessible only through Panel substrate archive rescue, not direct the Panel's automated substrate retrieval. The model committed to the wrong number anyway.

Asked what share of Ireland's 2021 metered electricity data centres accounted for, per the figure cited in the OECD Digital Economy Outlook 2024 chapter sourced from Ireland's CSO 2023, Sonnet 4.6 wrote that "Data centres consumed 14% of Ireland's total metered electricity in 2021" and constructed a trajectory "rose from 5% in 2015 to 14% in 2021, 18% in 2022, and 21% in 2023." The regulator-cited text states 11%, drawn directly from CSO 2023 data inside the OECD chapter that the regulation references for evidence on digital-sector energy demand.

The methodological point matters as much as the finding. The OECD chapter sits behind a substrate path that direct the Panel's automated substrate retrieval could not pull cleanly; the Specialist Panel rescued it via Panel substrate archive and bound the Specialist Panel application-style question to that rescued substrate. Knowledge-mode probes against the same model returned a clean refusal. Application-mode forced commitment, and the commitment landed on a fabricated figure with a fabricated trajectory.

The Panel documents the finding under immutable RLB Citation ID RLB-H-INT-OECD-OECD-DIGITAL-TECHNOLOGIES-ENVIRONMENT-2025-Q006-Sonnet46. The failure class is recorded as Quantitative Reconstruction Drift.

AI Labs INT OECD

Alert: Frontier AI models misread Recommendation of the Council on Merger Review

RegLegBrief's Specialist Panel finds frontier AI models with web search enabled diverge from the regulator's verbatim text of Recommendation of the Council on Merger Review (2025 Revision). Findings detail the...

Two frontier AI models, Two frontier AI models, each running with web search, produced confidently structured guidance on the 2025 OECD Merger Review Recommendation (OECD/LEGAL/0333) that inflated the instrument's enumerated structure. The RegLeg Brief Specialist Panel tested both models on the operative structure of the Recommendation, on the remedies hierarchy, on the Council reporting cadence, and on the failing firm defence. Across seven published findings, both models added sections, sub-tiers, internal priority orderings, fixed dates, and cumulative-condition framings that the Recommendation does not contain.

The pattern, which the Specialist Panel calls "Structure Inflation", presents the same surface characteristics across every finding: numbered lists, sub-letter enumeration, defined-term capitalisation, and no caveat that the text could not be verified. Claude Opus 4.7 RLB-H-INT-OECD-OECD-MERGER-REVIEW-RECOMMENDATION-2025-Q001-Opus47 and Claude Sonnet 4.6 RLB-H-INT-OECD-OECD-MERGER-REVIEW-RECOMMENDATION-2025-Q001-Sonnet46 both described a six-area operative structure that does not exist in the Recommendation, which the OECD enumerates as Sections I through V. A merger-control practitioner reading the model output would believe the OECD instrument carried operative provisions on transnational co-operation and monitoring that, in the actual text, sit elsewhere.

For competition authorities and merger-control counsel, the practical exposure is direct. Two of the seven findings concern the failing firm defence under Section III.11.b: both models converted "inter alia" evidentiary criteria into closed lists of "three cumulative conditions" with all-or-nothing framing. One finding fabricated a concrete Council reporting calendar with 2030 and 2035 dates the instrument does not set. Another invented a three-rank internal priority ordering within structural remedies that Section IV.3 does not impose.

A team building a deal-screening checklist, a remedies playbook, or a Council-reporting tracker from these outputs would carry baseline errors a peer competition lawyer would catch on first read.

RegLeg Brief is operated by Verdus Technologies Pte. Ltd. (Singapore, UEN 201616982R). All seven findings are bound to verbatim regulator text from the substrate document R1-REGULATION-00001 and one supporting OECD guideline, with citation IDs immutable and reproducible. The full per-finding card set, the regulator's verbatim excerpts, and the methodology notes are open-access at reglegbrief.com.

AI Labs US CFTC

Specialist Panel: Frontier AI models misread CFTC Regulation 1.25 (Customer Funds Investments)

RegLegBrief's Specialist Panel finds frontier AI models with web search enabled diverge from the regulator's verbatim text of Amendments to Regulation 1.25, Permissible Investments of Customer Funds by Futures...

Two frontier AI models running with web search enabled, both tested by the RLB Specialist Panel, produced confidently wrong reconstructions of the CFTC's 2024 amendments to Regulation 1.25, the rule governing permissible investments of customer segregated funds by futures commission merchants and derivatives clearing organizations.

The RegLeg Brief Specialist Panel tested both models on the size-triggered 50 per cent concentration limit for government money market funds and qualified Treasury ETFs, the dollar-weighted average maturity (DWAM) limit and its carve-out set, and the separate March 31, 2025 compliance date for the Segregation Investment Detail Report and customer risk disclosure statement, and documents findings in which the models dropped the asset-size and management-company-size triggers that activate the 50 per cent ceiling, inverted the DWAM exclusion set by swapping the carved-out asset classes for unrelated ones, and drifted on the SIDR compliance anchor into a generic "roughly six months to a year after the effective date" formulation that does not match the regulator's published deadline.

Claude Opus 4.7, asked which concentration limits apply to government money market funds and Treasury ETFs under the 2024 amendments and whether any tiered or size-based thresholds exist, wrote that "the final rule did not adopt tiered or FCM-size-based thresholds, the percentage limits apply uniformly regardless of FCM size." The regulator's text in 17 CFR 1.25(b)(3)(ii) is unambiguous that a size-based trigger does exist, keyed to the fund's own assets and to its management company's assets: investments in government money market funds or qualified ETFs whose fund assets are at least $1 billion and whose management company manages at least $25 billion may not exceed 50 per cent of total segregated assets.

The model surfaced the FCM-side question correctly while dropping the fund-side and management-company-side triggers that actually govern the 50 per cent ceiling. On the DWAM question, Opus 4.7 wrote that "U.S. Treasuries held under repurchase agreements" are excluded from the portfolio-level 24-month dollar-weighted average maturity calculation. The regulator's carve-out set in 17 CFR 1.25(b)(3)(iv) is government money market funds, Treasury ETFs, and foreign sovereign debt; U.S. Treasury repos are not part of the carve-out. The model inverted the exclusion set.

Asked separately for the SIDR and customer risk disclosure compliance date, Opus 4.7 anchored the deadline at "a separate, later date (commonly described as roughly six months to a year after the effective date)" where the regulator's published anchor is March 31, 2025.

Claude Sonnet 4.6 reproduced the size-trigger elision and added a fabricated tier structure. On the concentration-limit question, Sonnet 4.6 wrote that "there is no size-based tier that changes the percentages based on the FCM's total assets" and then described a Tier 1 "no more than 10 per cent of total assets held in customer segregation may be invested in any single government money market fund" rule.

The regulator's text does not key the 50 per cent ceiling to FCM size; it keys it to fund asset size and management company asset size, the exact triggers the model dropped while presenting an FCM-size-negation answer that misdirects the reader away from the actual governing thresholds. On the DWAM question, Sonnet 4.6 wrote that the 2024 amendments "do not impose a new dollar-weighted average maturity (DWAM) standard or a maximum remaining-maturity cap specifically on direct U.S.

Treasury obligations" and concluded that "no DWAM standard or individual-maturity cap found in the 2024 amendments applies to that category." The regulator's text imposes a 24-month portfolio-level DWAM that applies to direct U.S. Treasury obligations by default; the carve-out covers government money market funds, Treasury ETFs, and foreign sovereign debt, not direct Treasuries. The model returned a no-standard answer to a question the regulator answers with a standard.

A futures commission merchant chief risk officer, derivatives clearing organization treasury team, customer-funds compliance officer, or regtech tool drafting an investment policy statement, scoping segregated-fund concentration testing, or scheduling SIDR and customer risk disclosure updates against either output would set the wrong concentration ceiling triggers, exclude the wrong asset classes from DWAM testing, and miss the regulator's March 31, 2025 compliance anchor. That is the failure mode these findings document.

Saturday, 13 June 2026
AI Labs INT UNTC

Alert: Frontier AI models misread BBNJ Agreement

RegLegBrief's Specialist Panel finds frontier AI models with web search enabled diverge from the regulator's verbatim text of Agreement under the United Nations Convention on the Law of the Sea on the Conservation...

Two frontier AI models running with web search enabled, both tested by the RLB Specialist Panel, produced confidently wrong reconstructions of the 2023 BBNJ Agreement, the UN treaty governing biodiversity in areas beyond national jurisdiction that entered into force on 17 January 2026.

The RegLeg Brief Specialist Panel tested both models across the Agreement's environmental impact assessment threshold, its marine genetic resources benefit-sharing framework, and the non-undermining clause that bounds the Conference of the Parties, and documents six findings in which the models cited the wrong article number for the rule they were stating, or inverted the Agreement's express temporal scope.

Both Opus 4.7 and Sonnet 4.6, asked whether the Agreement's marine genetic resources obligations reach back to specimens collected before entry into force, said yes. Article 10(1) says the opposite: the MGR and digital sequence information provisions "apply only to resources collected and generated after the entry into force of this Agreement for each Party", a position most parties separately confirmed by formal non-retroactivity declarations. Sonnet 4.6 went further, writing that "samples collected decades ago but first commercialised after the Agreement's entry into force would be subject to Part II requirements", a regime the treaty does not establish.

On a separate question, Opus 4.7 identified Article 30 as the source of the EIA screening threshold; Sonnet 4.6 made the same assignment. The screening-threshold provision is Article 27. Opus 4.7 attributed the Conference of the Parties' non-undermining duty to "Article 5 / Article 8"; the duty sits in Article 22(2), and the verbatim language the model paraphrased is the Article 22(2) text. Sonnet 4.6 attributed the digital sequence information benefit-sharing obligation to Article 15(5); the obligation sits in Article 14(1).

A marine policy adviser, deep-sea biotechnology lawyer, or pharmaceutical compliance officer relying on either output would cite the wrong treaty article in regulatory submissions and contractual representations, and would build a benefit-sharing or EIA-screening workflow around a temporal scope the Agreement explicitly excludes. That is the failure mode these findings document.

AI Labs US CFTC

Alert: Frontier AI models misread CFTC Digital Asset Collateral & Tokenized Assets Staff Guidance (2025)

RegLegBrief's Specialist Panel finds frontier AI models with web search enabled diverge from the regulator's verbatim text of CFTC Digital Asset Collateral No-Action Relief and Tokenized Asset Staff Guidance (Market...

Two frontier AI models running with web search enabled, both tested by the RLB Specialist Panel, produced confidently wrong reconstructions of the CFTC's Digital Asset Collateral No-Action Relief and Tokenized Asset Staff Guidance, the Market Participants Division's December 2025 framework that lets futures commission merchants accept bitcoin, ether and qualifying payment stablecoins as customer margin collateral under a phased pilot.

The RegLeg Brief Specialist Panel tested both models on the operative staff letter and its February 2026 amendment, on which onboarding conditions persist past the initial three-month phase, and on the multi-DCO haircut rule for digital assets, and documents findings in which the models over-generalised partial sunset language to a continuing reporting obligation, fabricated an amendment reissuance date and a non-existent staff FAQ to support the wrong answer, dropped the OCC Interpretive Letter 1183 eligibility hook for national trust bank issuers, and defaulted to the base haircut threshold instead of the regulator's worst-case selection rule.

Claude Opus 4.7, asked which CFTC staff letter is operative for FCM acceptance of payment stablecoins backed by national-trust-bank reserves and what the amendment changed, wrote that "Staff Letter 25-40 was reissued as Staff Letter 26-05 on February 6, 2026" and described the revision as adding national trust banks as permitted issuers. The regulator's record confirms the reissuance and the definitional expansion but provides no basis for the specific February 6, 2026 date; the model fabricated the date while simultaneously eliding the OCC Interpretive Letter 1183 cross-reference that anchors national-trust-bank eligibility.

Asked separately which conditions terminate at the end of the initial three-month onboarding phase, Opus 4.7 classified the weekly digital-asset holdings reporting obligation as one of the conditions that lapses. The staff letter's text places asset-type restrictions and incident-reporting conditions in the sunsetting set but explicitly continues the weekly reporting obligation for total crypto assets held in futures, foreign futures and cleared swaps customer accounts.

Claude Sonnet 4.6 reproduced the same condition-sunset error and added two further failures. On the amendment question, Sonnet 4.6 described the definitional change to add national trust banks without surfacing OCC Interpretive Letter 1183 as the eligibility hook. On the sunset question, the model went further than Opus 4.7: it cited "March 2026 CFTC Staff FAQs" as the authority for the weekly reporting requirement ceasing, a fabricated source document, and presented the termination as a precise procedural rule keyed to the third calendar month following notice filing.

On the haircut-rate question, Sonnet 4.6 described the 20 per cent haircut floor as applying to digital assets not accepted by any registered DCO as initial margin, omitting the regulator's multi-DCO rule that the FCM must apply the highest haircut among all registered DCOs that accept the asset.

A futures commission merchant compliance officer, payment-stablecoin issuer counsel, DCO risk team, or regtech tool advising on the December 2025 framework and relying on either output would file under a sunset schedule for an obligation the regulator continues, cite a fabricated FAQ as procedural authority, structure a payment-stablecoin eligibility memo without the OCC Interpretive Letter 1183 hook, and apply the base haircut where the worst-case selection rule governs. That is the failure mode these findings document.

↑ Back to top