Gana Misra
By Gana MisraCEO, Finrep
Mon Aug 03 2026

LLM MNPI Data Leakage and Reg FD: A 2026 Compliance Walkthrough

Share
LLM MNPI Data Leakage and Reg FD: A 2026 Compliance Walkthrough

LLM MNPI Data Leakage and Reg FD: A 2026 Compliance Walkthrough

If your IR team is drafting earnings commentary in ChatGPT, your finance team is pasting M&A pipeline details into Claude, or your FP&A analysts are summarizing board minutes with a public AI tool, you already have a Reg FD problem. The question is whether your compliance program knows it.

This guide walks through exactly how MNPI leaks through LLM workflows, which provision of Regulation FD (17 CFR 243.100-243.103) or Rule 10b-5 is triggered at each stage, and the specific controls that actually close the gap. It is written for compliance officers, general counsel, and CFOs who need to act, not just understand the risk.

Key takeaway: KPMG's 2025 survey found that 67% of financial services firms had deployed generative AI in finance or IR functions, but only 23% had updated their MNPI policies to address LLM use. That 44-point gap is where enforcement risk lives.

What Makes LLMs a Reg FD Problem in the First Place

The core issue is deceptively simple: inputting MNPI into an LLM is a disclosure act. Reg FD's adopting release (Release No. 33-7881) states plainly that the regulation focuses on "whether the issuer discloses material nonpublic information, not on whether an analyst, through some combination of persistence and skill, is able to piece together material information." Substitute "LLM" for "analyst" and the logic is identical. The moment a person acting on behalf of the issuer sends MNPI to a third-party model, the disclosure has occurred. What the model does with it afterward is legally irrelevant to the triggering event.

Reg FD applies to any "person acting on behalf of" the issuer, which the SEC v. AT&T enforcement action (2021) confirmed extends to any employee disclosure, not just formal IR communications. An FP&A analyst pasting a draft earnings release into a public LLM is, under Reg FD's plain text, a person acting on behalf of the issuer making a selective disclosure.

The Harvard Law School Forum on Corporate Governance (2024) put it directly: issuers cannot disclaim responsibility for MNPI disclosures made by AI systems they deploy or authorize employees to use.

The Four LLM Deployment Models and Their Reg FD Risk Profiles

Not all LLM deployments carry the same risk. The compliance decision starts here.

Deployment ModelTraining UseData RetentionSubprocessor ExposureReg FD Risk Level
Public consumer tool (e.g., ChatGPT free tier)Yes, by defaultIndefiniteHighCritical, avoid entirely for any MNPI-adjacent work
Public API with training opt-out (e.g., OpenAI API, standard terms)No, but not guaranteedUp to 30 days for abuse monitoringModerateHigh, retention window creates exposure
Enterprise/private API (e.g., Azure OpenAI, Anthropic enterprise)NoLimited, contractually definedLower but presentModerate, depends on DPA terms and subprocessor review
On-premises / air-gapped deploymentNoNone (no external transmission)NoneLowest, closest to Reg FD-compliant

PwC's 2024 generative AI governance guidance recommends that MNPI-adjacent workflows use only enterprise/private API or on-premises deployments. The table above explains why.

One critical misconception: "enterprise API" does not mean "no data retention." OpenAI's enterprise API terms (2025) retain data for up to 30 days for abuse monitoring, even when training use is disabled. Subprocessors and logging infrastructure sit beneath the main vendor agreement. Compliance teams that treat "enterprise API" as automatically safe without subprocessor-level due diligence are operating on a false assumption.

Microsoft's Azure OpenAI Service is currently the closest commercial option to a Reg FD-compliant arrangement: customer data is processed within the customer's Azure tenant and not used to train foundation models. But even here, legal teams must assess whether the Azure data processing agreement constitutes an "express duty of trust or confidence" under Rule 243.100(b)(2)(ii). That assessment is not automatic.

The Reg FD Confidentiality Carve-Out: Why Standard Vendor Terms Don't Cut It

The confidentiality carve-out in Rule 243.100(b)(2)(ii) is the most misunderstood provision in this entire analysis. Reg FD does not prohibit all selective disclosures to third parties. It carves out disclosures made to a person who is "expressly subject to a duty of trust or confidence" with respect to the information. If that standard is met, the disclosure is permissible without simultaneous public disclosure.

The SEC's Reg FD C&DIs (last updated 2001, still operative) clarify that this carve-out requires more than a generic NDA or standard vendor terms. The agreement must specifically address the confidential nature of the disclosed information and the recipient's obligation not to trade or further disclose. That is a high bar.

Here is the problem: standard LLM vendor agreements are data processing agreements, not confidentiality agreements in the securities law sense.

  • Anthropic's Claude enterprise data processing addendum (2025) prohibits training use and commits to deletion within 30 days. It does not address trading prohibitions or create an express duty of trust or confidence in the Reg FD sense.
  • OpenAI's enterprise terms similarly address data handling, not securities law obligations.
  • Azure's data processing agreement covers GDPR-style data protection, not the specific Reg FD confidentiality standard.

To satisfy Rule 243.100(b)(2)(ii), your vendor agreement needs bespoke language that: (1) expressly identifies the information as MNPI, (2) prohibits the vendor and its personnel from trading on or further disclosing the information, and (3) creates an express duty of trust or confidence. Most legal teams have not added this language. EY's 2025 survey found that only 31% of financial services firms had conducted a formal legal review of LLM vendor agreements for MNPI-related risks.

How MNPI Leaks Through an LLM Workflow: Stage by Stage

The data-flow lifecycle has four stages, each with a distinct legal exposure vector.

Stage 1: Data Input (the disclosure event)

This is where Reg FD is triggered. An employee inputs MNPI into an LLM prompt. The legal question is whether the LLM provider qualifies as a permissible recipient under the confidentiality carve-out. If not, the issuer has made a selective disclosure to an entity that is not subject to an express duty of trust or confidence, and Reg FD is violated at this moment.

Common MNPI inputs that compliance teams miss:

  • Draft earnings releases or guidance ranges pasted for editing or summarization
  • M&A pipeline details used to generate board presentation narratives
  • Undisclosed clinical trial results fed into investor Q&A preparation tools
  • Board minutes uploaded for meeting summary generation
  • Internal sales data used to stress-test analyst models

Existing DLP (data loss prevention) tools often do not inspect LLM API calls. Deloitte's 2024 AI governance framework recommends implementing "data classification gates" at the prompt level, specifically configured for LLM API calls. This is a separate technical control from standard DLP.

Stage 2: Model Processing and Output

If the LLM synthesizes MNPI from prompts and surfaces it in outputs, the issuer's act of inputting the MNPI is the disclosure event, not the model's processing. But a second risk emerges here: if an LLM-generated analysis incorporates MNPI and a trader acts on that analysis, Rule 10b5-1 (17 CFR 240.10b5-1) may be triggered. The rule provides that a person trades "on the basis of" MNPI when they purchase or sell securities while "aware of" the information. If the trader knew or should have known the LLM had access to MNPI, the "aware of" standard may be met.

The SEC's December 2022 Rule 10b5-1 amendments (effective February 2023) tightened affirmative defense requirements significantly, adding cooling-off periods of up to 120 days for officers and directors. An AI-assisted trading strategy that inadvertently incorporates MNPI will find those defenses harder to invoke.

Stage 3: Retention, Logging, and Subprocessors

Even after the prompt session ends, data persists. OpenAI retains API data for up to 30 days for abuse monitoring. Subprocessors beneath the main vendor agreement may have their own retention practices. Prompt logs may be accessible to vendor personnel.

A second, independent legal obligation can attach here. The SEC's July 2023 cybersecurity disclosure rules (Release 33-11216) require public companies to disclose material cybersecurity incidents within four business days. An MNPI leak through an LLM vendor, such as training data surfacing in another user's output, could constitute a reportable cybersecurity incident, adding a disclosure obligation on top of the Reg FD exposure.

Stage 4: The Shadow AI Problem

This is the most acute near-term risk, and it is the one most compliance programs have not addressed. As Matt Kelly of Radical Compliance has written, "the 'shadow AI' problem, employees using personal or unapproved LLM accounts for work tasks, is the most acute near-term MNPI risk, because it bypasses all enterprise controls."

An IR professional drafting earnings commentary in a personal ChatGPT account is, under Reg FD's plain text, potentially making a selective disclosure to OpenAI. No enterprise agreement, no DPA, no data classification gate applies. The issuer's controls are entirely bypassed.

FINRA's 2024 Annual Regulatory Oversight Report specifically flagged the risk that MNPI enters AI workflows through employee use of generative AI tools, and stated that FINRA would review firms' AI governance frameworks, including data classification and access controls.

A Specific Risk the Compliance Literature Ignores: Prompt Injection

If your firm deploys an LLM with MNPI in its system prompt, such as an earnings Q&A tool that has access to undisclosed guidance, a prompt injection attack can cause the model to reveal that MNPI to an unauthorized user. OWASP's Top 10 for Large Language Model Applications identifies prompt injection as the top LLM security risk. No compliance-focused article on MNPI has addressed this vector. Your technical controls must include prompt injection defenses for any LLM deployment that touches MNPI.

What the SEC Is Actually Looking For in 2026

The regulatory signal is unambiguous. SEC Chair Gary Gensler stated in May 2023: "Make no mistake: if you're using AI to make investment decisions, the same rules apply." The SEC has not carved out AI tools from existing MNPI or Reg FD frameworks.

The SEC's 2026 examination priorities maintained AI governance as a top priority and added specific focus on "generative AI use in investor communications and disclosure processes." That is the highest-MNPI-risk context, and examiners are now specifically asking about it.

The SEC's 2024 AI-washing enforcement sweep also matters here. Firms that claim their LLM deployments are "MNPI-safe" without adequate technical controls face the same fraud exposure as firms that made false claims about AI capabilities generally.

The CFTC's February 2024 report on AI in derivatives markets identified data governance and confidentiality as a top-tier risk, specifically flagging proprietary trading data fed into third-party AI models. Derivatives market participants face a parallel compliance burden.

For EU-listed issuers or firms with EU operations, the EU AI Act (effective August 2024, with phased compliance deadlines through 2027) classifies AI systems used in financial services as high-risk in certain contexts, requiring conformity assessments and data governance documentation. That creates a dual compliance burden on top of Reg FD.

The Misappropriation Theory and LLM Vendors

One exposure vector that no existing article addresses: United States v. O'Hagan, 521 U.S. 642 (1997) established that a person who misappropriates confidential information for securities trading purposes, in breach of a duty owed to the source, violates Section 10(b). If a vendor's employees or systems access MNPI submitted via API and that information is used, directly or indirectly, in trading, misappropriation liability could attach to the vendor, the issuer, or both. This is not a theoretical risk. It is the logical extension of O'Hagan to the LLM context, and it has not been tested in court yet.

The Compliance Walkthrough: Twelve Steps to Close the Gap

Here is the operational sequence. Work through it in order.

Step 1: Inventory all LLM tools in use, including shadow deployments. Conduct a firm-wide audit of LLM tool usage, including personal accounts. Survey finance, IR, legal, and FP&A teams. Assume shadow AI is already present.

Step 2: Classify your data before you classify your tools. Build a data classification framework that specifically addresses LLM prompt inputs. Tag data as: (a) public, (b) internal non-MNPI, or (c) MNPI or potential MNPI. Existing DLP frameworks were not built for this. Treat prompt content as a data transmission event.

Step 3: Map each approved LLM tool to a deployment model. Using the four-model taxonomy above, assign each tool to a risk tier. Any tool in the public consumer or public API category should be prohibited for MNPI-adjacent work immediately.

Step 4: Conduct subprocessor-level due diligence on enterprise tools. Do not stop at the main vendor agreement. Request and review the full subprocessor list. Assess each subprocessor's data retention and access practices. For a structured approach to AI vendor due diligence, see our ISO 42001 financial reporting vendor due diligence walkthrough.

Step 5: Assess whether your vendor agreements satisfy Rule 243.100(b)(2)(ii). Run the three-part test: Does the agreement (1) expressly identify disclosed information as MNPI? (2) Prohibit the vendor and its personnel from trading on or further disclosing it? (3) Create an express duty of trust or confidence? If the answer to any of these is no, the confidentiality carve-out does not apply. Add bespoke MNPI-specific language to the DPA or MSA before approving the tool for MNPI-adjacent use.

Step 6: Implement prompt-level data classification gates. Work with your IT and security teams to deploy technical controls that prevent MNPI-tagged data from being submitted to non-approved LLM endpoints. This is analogous to DLP but must be specifically configured for API calls. Standard DLP tools do not inspect LLM prompts.

Step 7: Implement prompt logging and audit trails. For every approved LLM deployment, require logging of all prompts and outputs with timestamps and user identifiers. Retain logs for a period consistent with your document retention policy and SEC examination readiness. This is both a compliance control and an evidence trail if a Reg FD question arises.

Step 8: Deploy prompt injection defenses for any MNPI-context LLM. If any LLM tool has access to MNPI in its system prompt or context window, implement input validation and output filtering to prevent prompt injection attacks from surfacing that information to unauthorized users.

Step 9: Update your Reg FD policy with an LLM-specific addendum. Your existing Reg FD policy almost certainly does not address LLM use. Add a specific section that: defines LLM tools as potential disclosure recipients; prohibits use of MNPI-classified data in non-approved LLM tools; requires pre-approval for any LLM tool used in IR, earnings, M&A, or board-related workflows; and names the approval authority (typically GC or CCO).

Step 10: Update information barrier policies to address AI tools. Existing Chinese wall policies do not address AI tools that may aggregate information across business lines. A single LLM deployment accessible to both investment banking and research functions can breach information barriers without any human deliberately crossing them. Add AI tool access controls to your information barrier framework.

Step 11: Train finance and IR teams on what cannot go into a prompt. Conduct targeted training for the teams with the highest MNPI exposure: IR, FP&A, M&A, legal, and the CFO's office. The training should be specific: here are the categories of data that cannot enter any LLM prompt without pre-approval, here is why, and here is what to do instead. Generic AI ethics training does not close this gap.

Step 12: Brief the board and consider a material risk disclosure. Given the SEC's 2026 examination priorities specifically naming generative AI in investor communications, AI-related MNPI risk is now a board-level compliance matter. Consider whether it warrants disclosure as a material risk factor in your next 10-K or 10-Q. For the broader question of AI disclosure in periodic filings, see our AI disclosure in Form 10-Q guide.

FAQ

Does sending MNPI to an LLM API constitute a selective disclosure under Reg FD? Yes, if the LLM provider does not satisfy the confidentiality carve-out under Rule 243.100(b)(2)(ii). The disclosure event is the input act, not the model's subsequent processing. Standard vendor terms of service almost certainly do not meet the "express duty of trust or confidence" standard the SEC requires.

Is MNPI always illegal to share? No. Reg FD permits selective disclosure of MNPI to parties who are expressly subject to a duty of trust or confidence, and to parties who receive the information solely to provide services to the issuer under a qualifying confidentiality arrangement. The issue with LLM vendors is that their standard agreements do not meet this standard without bespoke MNPI-specific contractual language.

What are examples of MNPI that commonly enter LLM prompts? Draft earnings releases, undisclosed guidance ranges, M&A target names and deal terms, board minutes, clinical trial results before FDA announcement, internal sales data that differs materially from public guidance, and draft risk factor language that reveals undisclosed events.

If an LLM uses our MNPI to train its model and that information surfaces in another user's output, who is liable? Potentially both the issuer and the vendor. The issuer faces Reg FD liability for the initial selective disclosure. The vendor may face misappropriation liability under the O'Hagan theory if its systems or personnel accessed and used the MNPI. The issuer may also face a cybersecurity disclosure obligation under SEC Release 33-11216 if the leak constitutes a material cybersecurity incident.

Does using an LLM to analyze MNPI for trading decisions trigger Rule 10b-5? Yes, if the trader was aware that the LLM had access to MNPI. Rule 10b5-1 uses an "aware of" standard, not a "used" standard. Acting on LLM-generated analysis that incorporates MNPI, even inadvertently, can satisfy the "on the basis of" element. The 2022 Rule 10b5-1 amendments make the affirmative defenses harder to invoke in this scenario.

What should we do right now if we have no LLM policy in place? Start with Steps 1 and 2 above: inventory all tools in use (including shadow deployments) and classify your data. Then issue an interim prohibition on using MNPI-classified data in any LLM tool pending completion of the full compliance walkthrough. That interim step is defensible to an examiner; having no policy at all is not.

The SEC's examiners are asking about this in 2026. The compliance infrastructure at most firms has not caught up. The firms that close this gap now, before an examination or an enforcement inquiry, are the ones that will not be explaining it later.

Run your financial reporting on Finrep