Governance, compliance, and regulation of agentic AI
AI Act, NIST AI RMF, model risk management, and auditable evidence for agentic systems.
Learning objectives
- Place the real state of AI regulation applicable to trading agents as of July 2026, distinguishing rules in force, withdrawn proposals, and emerging practice in the US, the EU, and the UK.
- Explain what changed with SR 26-2 (Apr 17, 2026) relative to SR 11-7, why the explicit exclusion of generative and agentic AI creates a governance gap, and which principles of classical model risk management doctrine persist as examinable good practice.
- Reconstruct the AI washing enforcement cases (Delphia, Global Predictions, Rimar Capital) with firm, date, amount, and legal basis, and extract the doctrine: the Marketing Rule and fiduciary duty are already sufficient to sanction.
- Apply the judicial precedents on hallucinations (Moffatt v. Air Canada, Mata v. Avianca) and the 17a-4 record-keeping regime to agent-generated communications and traces.
- Design a fund's minimum governance artifacts: an AI use policy (data × model matrix), an agent inventory, immutable traces, a human-in-the-loop escalation matrix, and audit evidence.
- Justify, on regulatory grounds, why the patterns from modules 3 (HITL) and 9 (RiskGateway) are mandatory, not optional.
The previous modules built the architecture; this one answers the remaining question: who requires all of this, and what happens if it is missing. The answer, as of July 2026, is paradoxical: in the US, the classical model risk management framework has just deliberately excluded generative and agentic AI from its scope, while the SEC sanctions with the same old rules; in the EU, the AI Act has just postponed its high-risk regime by sixteen months, yet ESMA already requires AI to be included in the annual algorithmic trading self-assessment. There is a single common thread: responsibility cannot be delegated —not to the model, not to the vendor, not to the chatbot (BuilderWorld) .
10.1 The regulatory map in July 2026
10.1.1 SR 26-2: the rescission of SR 11-7 and the governance gap
For fifteen years, the Federal Reserve's SR 11-7 letter (Apr 4, 2011, with its parallel OCC 2011-12 and FDIC FIL-22-2017) was the operational reference for US model risk management: a broad definition of "model," three pillars (development; independent validation with conceptual soundness, ongoing monitoring, and outcomes analysis; governance with inventory, tiering, and effective challenge), and explicit coverage of third-party models —"we buy it, we don't build it" was never a defense (paperswithbacktest.com) .
On April 17, 2026, the Fed, OCC, and FDIC published the revised guidance —SR 26-2 / OCC Bulletin 2026-13— which replaces and rescinds SR 11-7 and SR 21-8 (Pinecone) . Material changes: a narrower definition of "model"; explicit exclusion of generative and agentic AI from scope; a full-relevance threshold of $30{,}000 million in assets; validation tied to materiality rather than annual calendars; a simplified regime for vendor models; and an express disclaimer —non-compliance with the guidance, by itself, does not generate supervisory criticism (Pinecone) .
The consequence is a documented governance gap: prompts containing client PII, MNPI, confidential terms, and proprietary strategies flow today to external providers outside any validation regime, outside any formal governance structure, and frequently outside the audit record; agentic systems that orchestrate traditional models sit squarely in that hole (From Concepts to Korean Language Benchmarks) . That the guidance is not enforceable per se does not amount to impunity: examiners still expect a coherent approach to AI risk, and downstream consequences remain in scope —if a generative component produces systematic errors that feed prices, signals, or expected-loss estimates, those outputs still belong to the model inventory (youngju.dev) . For an asset manager, moreover, SR 26-2 operates as a de facto standard in the due diligence of institutional investors and banking counterparties.
What persists from SR 11-7 as examinable good practice is its conceptual skeleton: a complete inventory of use cases (including shadow AI: Copilots and personal API keys are the most cited finding in exams (MyEngineerPath) ), documentation, independent validation with documented effective challenge, ongoing monitoring, and board accountability (paperswithbacktest.com) . The contrast with the UK is instructive: PRA SS1/23 (May 2023), in force, does explicitly cover AI —"including the management of the risks associated with the use of AI in modeling techniques such as machine learning"— with responsibility assigned to a named senior management function (SMF) (firecrawl.dev) . The same agent can fall outside the model risk perimeter in the US and inside it in the UK: a gray area that the multi-jurisdiction firm must resolve by design, not by omission (Pinecone) .
Table 10.1 — Regulatory map of AI in trading and asset management, by jurisdiction (as of Jul 29, 2026)
| Jurisdiction | Instrument | Status as of Jul 2026 | Implication for trading agents |
|---|---|---|---|
| US (banking) | SR 26-2 / OCC 2026-13 (Fed, OCC, FDIC) | In force since Apr 17, 2026; rescinds SR 11-7; excludes GenAI/agentic AI; $30{,}000 M threshold; not binding per se (Pinecone) | No mandatory validation regime for agents, but supervisory expectation of coherent governance; downstream outputs in scope (From Concepts to Korean Language Benchmarks) |
| US (securities) | SEC: Marketing Rule 206(4)-1, fiduciary duty, Reg BI; PDA conflicts rule | PDA rule withdrawn Jun 17, 2025; no final rule forthcoming (BirJob) ; 2026 exam priorities explicitly include AI (mSightFlow) | Ex post supervision: AI washing enforcement and disclosure–controls consistency exams (Apify) |
| US (broker-dealers) | FINRA Regulatory Notice 24-09 | In force since Jun 27, 2024; technology neutrality (Apify) | The entire rulebook (Supervision Rule 3110) applies to GenAI, including vendor tools (Apify) |
| US (records) | SEC Rule 17a-4 (amended Oct 2022) | In force; WORM or complete audit trail (OrbTop) | LLM outputs that are business communications —and their prompts— must be preserved; traces as audit trail (OrbTop) |
| EU | AI Act (Regulation 2024/1689) + Digital Omnibus | Bans and AI literacy in force; Art. 50 (transparency) from Aug 2, 2026; Annex III high risk postponed to Dec 2027 (Council, Jun 29, 2026) (Parse) | AI trading is not "high risk" today; transparency and training obligations do apply (Docs by LangChain) |
| EU (algorithmic trading) | MiFID II Art. 17 + RTS 6 (Regulation 2017/589); ESMA briefing Feb 26, 2026 | In force; AI must be integrated into the annual RTS 6 self-assessment (Docs by LangChain) | Full, non-delegable firm responsibility even when using third-party algorithms; retesting upon cumulative material changes (Docs by LangChain) |
| EU (ICT resilience) | DORA (Regulation 2022/2554) | Applicable since Jan 17, 2025 (RavenPack) | LLM/cloud providers supporting critical functions: register of information, audit rights, exit strategy (RavenPack) |
| UK | PRA SS1/23 (May 2023) | In force; explicitly includes AI; responsibility assigned to an SMF (firecrawl.dev) | Agents and LLMs used in modeling enter the model risk perimeter with a named owner (firecrawl.dev) |
| Singapore | MAS: FEAT (2018), Veritas, AI risk Guidelines | Guidelines in consultation (closed Jan 31, 2026; finalization expected in 2026) (nexusfi.com) | Full life-cycle governance, proportionality, GenAI-specific controls (hallucination, prompt injection, HITL) (nexusfi.com) |
The table condenses the asymmetry of the moment: the US deregulates in form (PDA withdrawn, SR 26-2 with a carve-out and a disclaimer) while supervising in practice through exams and enforcement; the EU regulates ex ante but postponed the hardest part of the AI Act, leaving the real weight on the MiFID II/RTS 6 and DORA sectoral anchors; the UK occupies the middle ground with a framework that does name AI. The operational reading: no jurisdiction exempts the governance system —where there is no validation rule there is enforcement by outcomes, and where the horizontal rule is postponed the sectoral one already requires integrating AI into the annual self-assessment. The convergence is the principle of non-delegable responsibility; the divergence lies only in the instrument (ex ante versus ex post), not in the outcome (BuilderWorld) .
In plain terms: a governance gap is not an amnesty
Imagine your city stops requiring structural inspection for a new construction material. Does that mean you can build however you want? No: if the building collapses, civil liability is intact; the insurer will still demand the old standard to underwrite a policy, and the bank won't finance you without it. That is exactly SR 26-2: it removes the obligation of the regime (formal validation of agentic AI), not the responsibility for outcomes.
Three concrete anchors survive the rescission, and they are worth keeping in mind every time someone says "it's no longer needed":
- The examiner expectation: supervisory teams still expect a coherent approach to AI risk, even though the guidance is no longer enforceable per se.
- Downstream outputs: if a generative component feeds prices, signals, or expected loss, those outputs remain in the model inventory.
- The de facto standard: in the due diligence of institutional investors and banking counterparties, SR 26-2 is used as a yardstick even though it does not bind.
10.1.2 Europe: the AI Act after the Digital Omnibus, ESMA, and the RTS 6 self-assessment
The AI Act (Regulation (EU) 2024/1689) is, as of July 2026, a calendar and not a snapshot. The prohibited practices of Art. 5 and the AI literacy obligation of Art. 4 (Feb 2025) are already enforceable, as are the general-purpose model obligations (Aug 2025). In four days —on August 2, 2026— the transparency obligations of Art. 50 enter into force: every chatbot that interacts with people (e.g., investor servicing) must disclose that it is AI, and synthetic content must be labeled, with penalties of up to 15 million euros or 3% of worldwide turnover (Parse) . The Digital Omnibus (adopted by the Council on Jun 29, 2026) postponed the Annex III high-risk regime to December 2027 and the Annex I regime to August 2028, but left Art. 50 intact (Parse) .
The most frequently confused point: the use of AI in algorithmic trading is not high risk today under the AI Act (Annex III does cover credit scoring, relevant for managers with a lending business), although ESMA warns that the classification is reviewed annually (Docs by LangChain) . Not being high risk is not being deregulated: the real European anchor is sectoral —Art. 17 of MiFID II and RTS 6 (Delegated Regulation (EU) 2017/589), with algorithm testing, annual self-assessment (Art. 9), stress testing, kill-switch, and pre-trade controls. ESMA's statement of May 30, 2024 set the doctrine: ultimate responsibility of the board and records documenting the use of AI —decision processes, data sources, algorithms, and any modification over time (BuilderWorld) . The supervisory briefing of Feb 26, 2026 closed the loop: the firm is fully and solely responsible even when using third-party algorithms; AI must be explicitly acknowledged in the annual RTS 6 self-assessment; compliance must understand how AI impacts the algorithm's decisions; and a "material change" that triggers retesting is any modification that could alter behavior, risk profile, or compliance posture, including cumulative changes —the most demanding clause for an agent whose model the vendor updates without notice (Docs by LangChain) . The Dutch AFM sums it up: algorithms —and agents— must be explainable, auditable, must not act in unintended ways, and must have clear lines of accountability (Docs by LangChain) .
Common pitfall: reading "not high risk" as "not regulated"
"AI trading is not high risk under the AI Act" is probably the most misread sentence of the module. The correct conclusion is not "we can relax until December 2027," but "the AI Act is not our main anchor — yet three layers do bite already":
- The AI Act, what's in force: AI literacy (Art. 4, since Feb 2025) and transparency (Art. 50, from Aug 2, 2026): a chatbot that talks to investors must disclose that it is AI, on pain of up to €15 M or 3% of worldwide turnover.
- The sectoral anchor: MiFID II Art. 17 and RTS 6 require testing, a kill-switch, pre-trade controls, and the annual self-assessment that must already acknowledge AI explicitly — with retesting upon cumulative material changes, including those the model provider makes without notice.
- DORA: if the LLM/cloud provider supports critical functions, you face the register of information, audit rights, and an exit strategy.
The AI Act is the horizontal layer; financial sector regulation is the vertical one. A postponement of the first suspends neither of the other two.
10.2 Enforcement and judicial precedents
10.2.1 The SEC's AI washing cases and 17a-4 record-keeping
With no AI-specific rule, the SEC found its path in the existing arsenal: Marketing Rule 206(4)-1, Sections 206(2) and 206(4) of the Advisers Act, and fiduciary duty. The doctrine was set in a single day —Mar 18, 2024, the first two AI washing actions in history— and hardened six months later with a fraud case. The 2026 exam priorities announce continuity: the Division will examine whether representations about AI capabilities are fair and accurate, whether controls are consistent with disclosures, and whether algorithms produce recommendations consistent with the stated strategies (mSightFlow) .
Table 10.2 — AI washing enforcement cases (SEC, 2024–2025)
| Firm | Date | Amount | Conduct | Legal basis |
|---|---|---|---|---|
| Delphia (USA) Inc. | Mar 18, 2024 | $225{,}000 | Claimed to use ML and client data in its investment algorithms without having done so (ADV, website, press releases, 2019–2023) (Apify) | Advisers Act 206(2), 206(4); Marketing Rule |
| Global Predictions Inc. | Mar 18, 2024 | $175{,}000 | Falsely declared itself the "first regulated AI financial advisor"; AI forecasts without evidence of the claimed model (Apify) | Advisers Act 206(2), 206(4); Marketing Rule |
| Rimar Capital USA / Liptz / Boro | Oct 10, 2024 | $310{,}000 in fines; $213{,}611 in disgorgement; 5-year bar for Liptz | Nonexistent "AI-driven" platform; ~$4{,}000{,}000 raised from 45 investors; inflated AUM ($16–20 M claimed vs <$2 M real) (Massive) | Antifraud; SEC order PR 2024-167 |
| Presto Automation Inc. | Jan 14, 2025 | No fine (cooperation and remediation) | "Presto Voice" presented as proprietary AI; >70% of "automated" orders required human agents (Github) | Securities Act 17(a)(2); Exchange Act 13(a) |
The amounts are modest, but the doctrine is the lesson: no new AI rule was needed —the same old rules on misleading statements sufficed— and the reputational and professional-bar effect far exceeds the fine. The conduct pattern is constant: claiming nonexistent AI capabilities, or third-party capabilities as one's own, without documentary evidence to back them. The direct internal consequence: every AI claim in an ADV, factsheet, or presentation must be traceable to a verifiable artifact (evaluations, traces, production records). Presto adds the listed-company nuance: AI washing also contaminates corporate reporting, and early cooperation mitigates the financial penalty but not the censure (Github) .
Worked case 1 — SEC v. Rimar Capital (Oct 10, 2024). Rimar Capital USA, its CEO Itai Liptz, and its director Clifford Boro raised roughly $4{,}000{,}000 from 45 investors by presenting an "AI-powered" automated trading platform that, according to the SEC's order, did not exist; they claimed AUM of $16–20 million against less than $2 million in reality, with misappropriation of funds. The settlement: Liptz paid $213{,}611 in disgorgement plus interest and a $250{,}000 civil penalty, with an associational bar and the right to reapply after five years; Boro paid $60{,}000 (Massive) . The case delimits the duty of technical truthfulness: if the investment memo says "autonomous agent with risk gating," the graph, the gating, and the execution evidence must exist —the difference between marketing and reality is, since 2024, a nominal enforcement risk, not a theoretical one.
The second leg is record-keeping. Rule 17a-4 (amended Oct 2022) requires electronic records to be kept in non-rewritable, non-erasable format (WORM) or through a complete audit trail that allows the original to be recreated, with typical retention of 3–6 years (OrbTop) . The off-channel communications sweep has accumulated since 2021 more than $2{,}000 million in fines: JPMorgan, $200{,}000{,}000 (Dec 2021); sixteen firms, $1{,}100{,}000{,}000 in a single day (Sep 27, 2022) (RavenPack) . The analogical extension is immediate: a "chat with the model" about business is a communications channel; if an LLM generates or transforms research notes, client messages, or memos, those outputs —and, in examination practice, the prompts and contexts that produced them— must be preserved (OrbTop) . An agent without a trace archive is, under this reading, the 2026 equivalent of the trader with a personal WhatsApp.
In the UK, the reference cost of failure is Citigroup Global Markets Ltd (May 2024): the PRA imposed £33{,}880{,}000 and the FCA £27{,}766{,}200 —a total of £61{,}600{,}000— for failures in trading systems and controls (2018–2022), crystallized on May 2, 2022 (a trader entered $444{,}000{,}000 instead of $58{,}000{,}000; $1{,}400 million were executed on European exchanges) (firecrawl.dev) . The transferable lesson: the absence of hard preventive blocks and the inadequate calibration of controls is a governance failure, not a technology one —the exact regulatory justification for the RiskGateway of module 9.
Worked example: auditing a factsheet with the Delphia doctrine
Suppose the distribution team proposes this paragraph for a factsheet: "Our proprietary AI engine generates signals with deep learning on alternative data, under continuous human supervision." Let's apply the test an examiner would run after Delphia and Global Predictions: every claim must be traceable to a verifiable artifact.
| Claim | The examiner's question | Artifact that backs it | Case that anchors it |
|---|---|---|---|
| "Proprietary AI engine" | Is it yours or a vendor's? | Contract + documented architecture; if it's a third-party API, you don't say "proprietary" | Global Predictions ("first regulated AI advisor") |
| "Generates signals with deep learning" | Which model, with what evidence? | Production traces, evaluations, version records | Delphia (claimed to use ML without having done so) |
| "On alternative data" | Which sources, under what licenses? | Dataset inventory and usage agreements | Delphia (client data not used) |
| "Continuous human supervision" | Who, when, with what record? | HITL checkpoints with approver and decision | Presto (>70% required humans) |
Three of the four sentences could be true and still be sanctionable if the artifact does not exist: the doctrine does not punish ambition, it punishes the gap between marketing and production without documentary evidence. And a nuance that is easy to forget: the chat with the corporate LLM that marketing used to draft that paragraph is itself a record under the expansive reading of 17a-4 — the factsheet and its conversational draft must be producible together.

The figure puts the module's amounts in perspective: AI washing fines are measured in hundreds of thousands of dollars, while record-keeping and trading-controls failures are paid in hundreds of millions. Don't read this as "AI washing is cheap": the real cost of Delphia or Rimar is reputational and professional-bar related, and the lesson is asymmetric — lying about AI is exposed by a document; failing to archive what the AI did is exposed by the entire exam.
10.2.2 Liability for hallucinations and MNPI in prompts
If AI washing punishes saying too much about AI, hallucination case law punishes what AI says too much of on the firm's behalf.
Worked case 2 — Moffatt v. Air Canada (2024 BCCRT 149, Feb 14, 2024). The airline's chatbot invented a retroactive bereavement refund policy that did not exist, and a passenger booked relying on it. Air Canada argued that the chatbot was "a separate legal entity responsible for its own actions"; the tribunal called the argument a "remarkable submission" and found the airline liable for negligent misrepresentation: CA$812{,}02 (BuilderWorld) . The amount is trivial; the doctrine is not: what the system says, the firm says. In asset management, quarterly letters, factsheets, DDQ responses, and market commentaries generated with an LLM are "communications" for purposes of the Marketing Rule, Art. 24 of MiFID II, and Art. 50 of the AI Act; European legal literature requires those contents to be explainable, auditable and consistent with the other channels (mediawatcher.ai) . The open question was posed in 2026 by Judge Philip Jeyaretnam (Singapore): can a deployer be exempted by demonstrating adequate training, testing, and monitoring, or through disclaimers? (Benzinga) . Everything points to documented "algorithmic due diligence" —traces, evaluations, human review— being the only available line of defense.
The second precedent, Mata v. Avianca (SDNY, Jun 22, 2023), sanctioned two lawyers with $5{,}000 for filing six nonexistent cases generated by ChatGPT —the model even "confirmed" its own fabricated citations (fitgap.com) . The technical lesson underpins the course's grounding pattern: an LLM cannot verify its own output; verification requires an authorized external source, and every cited figure must exist verbatim in it.
Table 10.3 — Hallucination and record-keeping precedents: case, jurisdiction, outcome, lesson
| Case | Jurisdiction | Outcome | Lesson for trading agents |
|---|---|---|---|
| Moffatt v. Air Canada (2024 BCCRT 149) | Canada (BC civil tribunal) | CA$812{,}02 for negligent misrepresentation; "separate entity" argument rejected (BuilderWorld) | The firm answers for every statement made by its chatbot/agent; client-facing outputs are regulated communications |
| Mata v. Avianca (Jun 22, 2023) | US (SDNY) | $5{,}000 sanction; 6 citations fabricated by ChatGPT (fitgap.com) | An LLM cannot verify itself; mandatory citation to an authorized source (RAG with grounding) |
| Off-channel communications sweep (2021–2024) | US (SEC/CFTC) | >$2{,}000 M accumulated; $1{,}100 M in a single day (Sep 27, 2022) (RavenPack) | Every business communication channel must be archived; the chat with the corporate LLM is no exception |
| Citi CGML (May 2024) | UK (PRA + FCA) | £61{,}600{,}000 combined for deficient trading controls (firecrawl.dev) | Hard pre-trade limits and documented calibration are governance, not optional engineering |
The joint reading delimits the deployer's liability perimeter: it answers for what the agent says (Air Canada), for what it cites (Mata), for what it leaves unrecorded (record-keeping), and for what it executes without limits (Citi). None of this requires bad faith: negligent misrepresentation, the unverified citation, and the miscalibrated control are enough. For a LangGraph agent in production, the architecture translates into obligations: grounding with verifiable citations, an immutable trace archive, hard blocks in the execution gateway, and signed human review of every external communication. Each column of the table corresponds to a control in the pipeline: source verification, record preservation, pre-trade limits, and demonstrable human supervision. Those who do not document will have no defense to offer (Benzinga) .
Risk note — MNPI in prompts: a potential per se violation. Entering material non-public information (earnings previews, M&A, positions) into an external AI tool can constitute by itself a violation of securities laws, even if no one trades on it (businessmodelcanvastemplate.com) . The case that fixed the industry's awareness was Samsung (Apr–May 2023): in twenty days, engineers pasted confidential source code and an internal meeting transcript into ChatGPT in three incidents; the company restricted use and accelerated an internal alternative (SmarterWay.AI) . Wall Street reacted before any other sector: JPMorgan restricted ChatGPT in early 2023, and Goldman Sachs, Citi, Bank of America, Deutsche Bank, and Wells Fargo followed (businessmodelcanvastemplate.com) . Information barriers must be extended to LLM instances: environments segregated by side of the wall, DLP that blocks the pasting of wall-crossed documents, restricted lists at the gateway's prompt layer, and logging of the system's own wall-crossing (who, when, what information, when it is "cleansed") (businessmodelcanvastemplate.com) .
In plain terms: what the system says, the firm says
The Air Canada case fits in one sentence: a customer asked the chatbot about bereavement refunds, the chatbot invented a generous policy, and when the airline refused to honor it, it argued that the chatbot was "a separate legal entity responsible for its own actions." The tribunal devoted two words to the argument —remarkable submission— and ruled against the airline: the company speaks through its bot.
Translation to asset management: every quarterly letter, DDQ response, or market commentary that leaves your pipeline carries your signature, even if the model wrote it. And this is where Mata v. Avianca fits: an LLM cannot verify its own output —in that case it even "confirmed" citations it had fabricated itself— so verification has to come from an authorized external source. That is why the course's grounding pattern is not aesthetics: it is the only architecture that turns "the model said it" into "the firm can prove where it came from." Disclaimers ("AI-generated content, may contain errors") have yet to save anyone; the only serious line of defense is documented algorithmic due diligence.
10.3 Practical governance: observability as a defense
10.3.1 From policy to the file: inventory, traces, and HITL
The SR 26-2 gap leaves firms without a rule telling them how to govern agentic AI, but fully responsible for the outcomes. In that hole, the industry has converged on a five-layer governance architecture that this module presents as emerging practice, not as regulation:
The first layer is the living inventory of use cases —of agents, not just models—: name, owner, provider, version, input data, intended use, and criticality, including shadow AI (MyEngineerPath) . The second is the AI use policy, whose core is a data × model matrix: public data to any LLM; sensitive internal data, only to enterprise instances with no-retention and no-training contracts; MNPI and proprietary strategies, only to segregated or on-prem models with information barriers; personal data, conditional on legal basis and minimization (SmarterWay.AI) . The third —the one that gives this section its title— is traces as evidence: to serve before an examiner they must be immutable (WORM or chained hashing, retention aligned to 17a-4 or MiFID II), bind identity and version (agent_id, model_id, model_version, system prompt, active tools), record decisions and not just outputs (alternatives considered, human approval points), and be retrievable in readable format within examination deadlines (OrbTop) . The fourth is graduated HITL: human-on-the-loop with sampling for internal research, human-in-the-loop with prior approval for client communications, human-over-the-loop with hard limits and a kill switch for execution —with override records as a KRI, because an approver who accepts 100% in two seconds is an exam finding, not a control (Docs by LangChain) . The fifth is audit: an annual RTS 6 self-assessment that integrates AI, record production under 17a-4, and quarterly board reporting (BuilderWorld) .
In institutional practice — how a fund documents its agents. A systematic fund with three agents in production (RAG research, daily risk monitor, quarterly commentary generator) implements the file today as follows: every run leaves a LangSmith trace with the complete tree —prompt, retrieved documents, tool calls with arguments, output, latencies, tokens— tagged with
agent_idandmodel_version; every HITL interruption persists a LangGraph checkpoint that freezes the graph state at the moment of approval, with the approver's name and decision. Nightly, a job exports traces and checkpoints to WORM storage with seven-year retention, indexed by date, agent, and ticker. When the examiner asks for "the May 14 decision on X," compliance replays the full trace in minutes: what the agent saw, what it computed, what it proposed, and who approved it. The practice maps the 17a-4 audit-trail route and paragraph 24 of the ESMA statement (BuilderWorld) : it is the market's answer to the SR 26-2 gap —emerging practice, not an explicit requirement of any rule— and, precisely for that reason, the only defense available against Judge Jeyaretnam's question (From Concepts to Korean Language Benchmarks) .
Two technical reasons make this discipline structural and not optional. First, non-determinism and vendor updates: pre-deployment validation cannot govern a system that changes after deployment —a structural incompatibility between the classical framework and the technology— so control shifts to continuous monitoring, version pinning, and revalidation triggers (Institutional Asset Manager) ; the RTS 6 definition of cumulative "material change" turns every silent vendor update into a potential retesting event (Docs by LangChain) . Second, hallucination as an inherent risk: not eliminable by architecture, so control is layered —grounding, mandatory citations, clean rejection, human review— and not "patching the model" (contextanalytics-ai.com) . If each review layer catches a fraction \(r\) of errors and there are \(n\) independent layers, the residual probability that a hallucination reaches the client is \(p_{\text{res}} = (1-r)^{n}\): the quantitative justification for why HITL must be real and distributed, not ceremonial and concentrated.
The cost of compliance must be budgeted with numbers. The reference parliamentary estimate puts the AI Act between €320{,}000 and €600{,}000 per entity under high-risk classification (written question E-001210/2026, Mar 24, 2026) (Apify) ; transparency implementations run around €15{,}000–31{,}000 in the first year, with 30–50% recurring annually in the most demanding regimes (Stock Market: From Zero to Hero) . The asymmetry matters: the fixed cost of validation and archiving weighs more on small managers and pushes toward concentration in large players and in the few vendors capable of signing DORA-style contracts —the third-party dependency vulnerability that the FSB flagged to the G20 in November 2024 (tessl.io) .
The conclusion closes Block IV: the patterns of modules 3 and 9 were not engineering preferences. The HITL interrupt exists because ESMA requires meaningful human oversight and ceremonial approval is an exam finding (BuilderWorld) ; the RiskGateway with hard limits exists because Citi paid £61{,}600{,}000 for not having them calibrated (firecrawl.dev) ; the trace archive exists because, after SR 26-2, observability is the only defense the firm can build for itself (From Concepts to Korean Language Benchmarks) . Those who build without observability build without a legal defense.
Common pitfall: ceremonial HITL
The most uncomfortable exam finding is not the absence of human review, but its caricature: an approver who accepts 100% of proposals with a median of two seconds per decision. That "human in the loop" does not add a control layer; it adds latency and a false sense of coverage. Worse still: it leaves a documentary trail that proves the opposite of what it intended.
How a firm detects it itself before the examiner does, with KRIs over the override log:
- Override rate ≈ 0% sustained (no real human always agrees with the model).
- Median review time incompatible with reading the proposal and its evidence.
- Approvals in bursts (50 decisions in 4 minutes) or outside working hours.
The link to the formula \(p_{\text{res}} = (1-r)^{n}\) is direct: a ceremonial layer has \(r \approx 0\) and \((1-0) = 1\) —it multiplies by one, it reduces nothing—. HITL only counts in the equation if it is real.
Worked example: how many review layers are needed
The quantitative justification of distributed HITL is best seen with numbers. Suppose independent review layers that each catch a fraction \(r = 0{,}8\) of the hallucinations that reach them. The residual probability that an error reaches the client is \(p_{\text{res}} = (1-r)^{n}\):
| Layers \(n\) | \(p_{\text{res}}\) | Marginal reduction |
|---|---|---|
| 1 | 20% | — |
| 2 | 4% | 16 p.p. |
| 3 | 0.8% | 3.2 p.p. |
| 4 | 0.16% | 0.64 p.p. |
Two design readings. First, returns diminish: going from one to two layers is the most profitable decision; the fourth layer barely moves the needle and does add latency and fixed cost —the compliance cost the module quantifies at €320,000–600,000 per entity under the high-risk regime—. Second, and more important: the formula demands independence. Two layers running the same model with the same prompt share failure modes —if the hallucination fools the first, it probably fools the second— and the effective \(r\) collapses. That is why layers must fail differently: grounding verification against the source, numeric cross-check with a deterministic tool, human review with evidence in view. And \(r\) is not assumed: it is measured by sampling against ground truth. Exercise 3 will ask you to repeat this calculation with other values of \(r\) and \(n\).
From execution to the file: anatomy of a defensible trace
The module describes the five governance layers; this diagram drops one level and shows the physical journey of the evidence from the moment the agent executes until the examiner receives an answer:
Notice that each link answers a specific requirement of the module: the LangSmith trace and the checkpoint bind identity and version (agent_id, model_version) in addition to the human decision; the WORM export with chained hashing satisfies the 17a-4 audit-trail route; and the replay "in minutes" is what the ESMA statement expects when it asks for records documenting the use of AI over time. If any link is missing, the whole file loses value: a trace without a checkpoint proves what the agent computed but not who approved it; a checkpoint without an immutable archive does not survive the question "how do I know you didn't edit it afterwards?".
Exercises
- AI use policy. Draft the AI use policy of a fund with $2{,}000 M in AUM and long/short strategies: a data × model matrix for five data classifications and three deployments (public API, enterprise instance with zero-retention, on-prem model), with treatment of LLM-instance wall-crossing and an exceptions procedure. Justify each cell with the rule or precedent that anchors it.
- Mapping to RTS 6. Map each component of the module 9 execution agent (LangGraph graph, RiskGateway, kill-switch, deterministic tools, traces) to the RTS 6 requirements: testing, annual self-assessment (Art. 9), pre-trade controls, real-time monitoring. Identify which life-cycle events (vendor version change, new tool, system prompt change) trigger retesting under the cumulative definition of "material change."
- HITL escalation matrix. Design the escalation matrix for three agents (internal research, quarterly investor commentary, rebalancing execution): autonomy by use case, quantitative escalation criteria (order size, benchmark deviation, grounding confidence score), override KRIs, and ceremonial-approval detection. Estimate \(p_{\text{res}} = (1-r)^{n}\) for \(r \in \{0{,}7, 0{,}9\}\) and \(n \in \{1, 2, 3\}\) and defend the proposed number of layers.
- Analysis of an enforcement case. A two-page report on SEC v. Rimar Capital (PR 2024-167, Oct 10, 2024): timeline, false statements, legal basis, sanctions per person, and doctrine for AI disclosures in ADV and marketing. Conclude with the three internal controls that would have prevented the case and how they would be evidenced before an examiner.
- Examination file. Simulate a request from the Division of Examinations under the 2026 priorities: list the artifacts the fund must produce to demonstrate that its AI representations are fair and accurate and consistent with its controls, specifying for each one the format, retention, and the 17a-4 requirement it satisfies.
- The SR 26-2 gap. A risk committee asks: "If SR 26-2 excludes agentic AI and is not binding, why budget for governance?" Prepare the CRO's answer in ten lines, citing the nature of the guidance, examiner expectations, downstream in-scope consequences, and the contrast with PRA SS1/23 for the London subsidiary.
Module 10 of 15 — LangChain for Quantitative Trading. Data and versions verified as of July 2026; prices and market figures subject to change. LangChain 1.3.14 / langchain-core 1.5.2.
Elite resources to go deeper
- SR 11-7 — Supervisory Guidance on Model Risk Management (Fed, 2011): the original text that SR 26-2 rescinds; you need to know it to understand what has been lost and what persists as good practice.
- SEC PR 2024-36 — charges against Delphia and Global Predictions: the first two AI washing actions in history, with the Marketing Rule doctrine applied to AI claims.
- SEC PR 2024-167 — SEC v. Rimar Capital: the module's worked case; the order details false statements, inflated AUM, and sanctions per person.
- FINRA Regulatory Notice 24-09: the technology-neutrality reference — the entire rulebook, including Supervision Rule 3110, applied to vendor GenAI.
- PRA SS1/23 — Model risk management principles for banks: the live UK contrast that does name AI and assigns responsibility to an SMF.
- European Commission — AI regulatory framework: the official page for tracking the AI Act calendar (transparency, high risk) after the Digital Omnibus.
- NIST AI Risk Management Framework: the common vocabulary (govern, map, measure, manage) that organizes any defensible AI use policy.
- MAS — FEAT principles: the Singapore reference on fairness, ethics, accountability, and transparency, the basis of its AI risk guidelines.
Check your understanding
Self-assessment with instant feedback. No scores are stored: it is just for you.