NLG MD&A vs Human-Drafted Accuracy: Where AI Fails SEC Scrutiny
Disclosure teams at mid-to-large public companies are using AI drafting tools, from Microsoft Copilot to Harvey to proprietary NLG systems, to produce first-draft MD&A at speed. The question CFOs and disclosure counsel are actually asking is not whether the numbers are right. It is whether the narrative will survive SEC review.
The answer depends on which kind of accuracy you mean. NLG-drafted MD&A tends to be numerically consistent. It also tends to be generically worded, light on forward-looking analysis, and structurally blind to several specific Item 303 requirements. And as of August 2026, the SEC is using AI to read your filings before a human attorney does.
Key takeaway: NLG output can be factually accurate and still fail SEC scrutiny. The specificity gap, the trend carry-forward blind spot, and the earnings call consistency problem are the three failure modes most likely to generate comment letters, and the SEC's new AI reviewer is calibrated to detect all three.
What the Research Says About NLG MD&A vs Human-Drafted Quality
NLG-generated MD&A is more internally consistent with financial data but measurably less specific, less forward-looking, and more similar to industry boilerplate than human-drafted equivalents. That is the core finding of Cao, Jiang, Yang, and Zhang (2020), the most-cited academic study on AI-generated MD&A disclosure quality. The researchers used readability metrics, specificity scores, and cosine similarity to peer filings to compare NLG output against human-drafted sections across a large sample of 10-K filings.
The specificity gap is the critical finding. NLG systems produce language that is semantically close to the average of what companies in the same industry say. That is, by definition, boilerplate. The SEC has flagged boilerplate as a disclosure failure since its 2003 MD&A interpretive release (Release No. 33-8350), which warned that "boilerplate disclosures and other generic language generally are not helpful and would detract from the purpose" of MD&A.
One important caveat: the Cao et al. study examined NLG tools from 2018 to 2019, which are far less capable than 2025 to 2026 large language models. Modern LLMs can produce more fluent, less obviously templated prose. But the structural limitation persists regardless of model capability: an AI system does not have management's perspective on forward-looking business conditions. It cannot know what the CFO said on the earnings call, what the board discussed about a deteriorating customer relationship, or why a margin trend is expected to reverse. That judgment gap is what the SEC's "eyes of management" standard is designed to require.
How the SEC's 2026 AI Reviewer Detects NLG Output
The SEC confirmed in August 2026 that it is using AI to "sift through vast quantities of public company filings and identify emerging risks and trends," in connection with the creation of the new Financial Reporting and Accounting Unit announced August 5, 2026. The unit is led by Timothy Zimmerman and reports to Osman Nawaz, Principal Deputy Director of the Division of Enforcement. For a full breakdown of the unit's mandate, see Finrep's SEC AI-powered review analysis.
What makes this consequential for disclosure teams is not just that the SEC is reading more filings. It is that the SEC's AI is specifically calibrated to detect the patterns NLG systems produce. Based on the SEC's public statements and its EDGAR infrastructure, the system appears to perform at least four functions:
- Consistency checking. NLP tools cross-reference quantitative claims in narrative text against XBRL-tagged financial data. An MD&A that states "revenue grew 12% in Q2" when XBRL-tagged actuals show 11.3% growth creates a flaggable inconsistency. NLG tools that generate narrative from rounded or approximated figures without checking against XBRL actuals create this risk systematically.
- Semantic boilerplate detection. LLM-based analysis assesses disclosure language for specificity relative to peer filings. A company cannot avoid flagging simply by paraphrasing generic language; the SEC's AI recognizes semantic equivalence across paraphrased boilerplate. This is the pattern NLG output most reliably produces.
- Filing history anomaly detection. AI systems compare specific metrics across multiple filing periods for a single company and flag patterns inconsistent with disclosed operating trends.
- Cross-company peer benchmarking. Statistical outliers in disclosed metrics versus industry peers are surfaced for human review.
The implication is direct: AI-drafted MD&A that would have survived keyword-based screening may not survive semantic pattern recognition. The SEC's reviewer is, in effect, trained on the same corpus of industry filings that NLG systems use as training data, and it is looking for the similarity that NLG output produces.
NLG vs Human MD&A: The Failure Mode Comparison
The table below maps the documented failure modes of NLG-drafted MD&A against the specific Item 303 requirements they violate and the SEC comment letter categories they trigger.
| NLG Failure Mode | Item 303 Requirement Violated | SEC Comment Category |
|---|---|---|
| Generic trend language, high cosine similarity to peer boilerplate | S-K 303(b): known material trends and uncertainties | Boilerplate; insufficient specificity |
| Paraphrasing cash flow statement instead of explaining drivers | S-K 303(b): liquidity and capital resources | Failure to explain primary drivers |
| Omitting "why useful" and "how management uses" for KPIs | S-K 303(b): key performance indicators | Incomplete KPI disclosure |
| No detection of KPI methodology changes period-over-period | S-K 303(b): KPI consistency | Failure to disclose methodology change |
| Drafting each period independently, dropping prior trend disclosures | Item 303 trend carry-forward obligation | Omission of required trend update |
| Generating narrative from financial data, not earnings call context | Consistency with other disclosures | Inconsistency with earnings release or transcript |
| Using pre-2021 tabular contractual obligations format | S-K 303(b)(1): material cash requirements | Deprecated disclosure format |
| Migrating forward-looking content from MD&A to footnotes | PSLRA safe harbor scope | Loss of safe harbor protection |
Each of these is a structural gap, not a model quality problem. A more capable LLM produces more fluent prose with the same underlying gaps, because the gaps arise from what the model does not have access to, not from how well it writes.
The Five Specific Item 303 Gaps NLG Tools Produce
1. The Specificity Gap and the Boilerplate Problem
Deloitte's Financial Reporting Manual (Topic 9) is explicit: registrants should not simply repeat items reported in the statement of cash flows. They should "focus on the primary drivers of and other material factors necessary to an understanding of the registrant's cash flows and the indicative value of historical cash flows." NLG tools that auto-populate the liquidity section by paraphrasing the cash flow statement produce exactly the output this guidance prohibits.
The same principle applies to results of operations. The SEC's Release No. 33-8350 requires MD&A to provide investors with the ability to see the company through the eyes of management. An NLG system describing a margin decline as "impacted by increased cost of goods sold" satisfies the letter of disclosure while providing none of the management perspective the standard requires.
2. The KPI Disclosure Gap
For any KPI or non-GAAP metric disclosed in MD&A, Deloitte's FRM specifies three required elements: a clear definition and calculation methodology, a statement of why the metric provides useful information to investors, and a statement of how management uses the metric. NLG tools that auto-populate KPI disclosures from prior-period templates frequently omit the second and third elements, producing a specific Item 303 compliance gap that SEC staff routinely flag.
When a KPI's calculation methodology changes period-over-period, the registrant must also disclose the differences in calculation, the reasons for the change, and the effects on amounts disclosed. An NLG system trained on prior-period filings will not detect a methodology change. That detection requires human judgment.
3. The Trend Carry-Forward Blind Spot
This is the most underappreciated NLG failure mode. Bass Berry's MD&A best practices guidance notes that when a trend is disclosed in one period's MD&A, the registrant may be required to continue disclosing and updating that trend in subsequent periods. It cannot simply be deleted. NLG tools that draft each period's MD&A independently, without reference to prior-period disclosures, will systematically fail to carry forward required trend disclosures.
The legal stakes are higher than a comment letter. As Bass Berry states directly: "The failure to disclose known trends can give rise to exposure from Rule 10b-5 allegations from private parties as well as SEC civil actions." A trend that appeared in your Q3 10-Q and disappeared from your Q4 10-K without explanation is a red flag for both the SEC's AI reviewer and plaintiffs' counsel.
4. The Earnings Call Consistency Problem
NLG tools generate MD&A narrative from financial data inputs. They do not read earnings call transcripts. This creates a structural inconsistency risk that is a documented SEC comment trigger. Bass Berry is explicit: "The SEC Staff will review not only earnings releases, but also earnings call transcripts, and any material inconsistencies between these disclosures and periodic report disclosures can be a topic of SEC comment."
If management discussed a specific customer concentration risk on the Q&A portion of the earnings call, that context should inform the MD&A. An NLG system drafting from financial tables alone will miss it. Human drafters, who participate in or review the earnings call, would not.
5. The PSLRA Safe Harbor Migration Risk
This is a legal risk that disclosure counsel using AI tools need to understand explicitly. The PSLRA safe harbor for forward-looking statements applies to MD&A disclosures accompanied by meaningful cautionary language. It does not apply to financial statement footnote disclosures. Bass Berry flags that NLG tools which migrate forward-looking content between MD&A and footnotes could inadvertently strip PSLRA protection from statements that would otherwise be covered.
This is not a hypothetical. NLG tools optimizing for disclosure completeness may move content to wherever a template has space, without awareness of the legal consequences of section placement.
The 2020 Item 303 Amendments: A Specific Template Risk
The SEC's Final Rule Release No. 33-10825, adopted November 19, 2020 and effective February 10, 2021, eliminated the prescriptive tabular contractual obligations disclosure and replaced it with a principles-based requirement to discuss material cash requirements from known contractual and other obligations. NLG tools trained on pre-2021 filing templates may still generate the deprecated tabular format.
This is a concrete, verifiable compliance gap. If your NLG tool's training data skews toward filings from 2018 to 2020, check the contractual obligations section output before filing. The 2026 compliance walkthrough on Finrep covers the full Item 303 modernization requirements in detail.
Can AI-Drafted MD&A Pass SEC Review With Human Oversight?
Yes, but only if the human review process is specifically designed to catch the gaps NLG tools produce. A general editorial review that checks for readability and factual accuracy is not sufficient. The review must be structured around the specific failure modes above.
For context on what the SEC actually requires from AI-assisted disclosures, see Finrep's analysis of SEC AI-generated MD&A requirements. There is no current SEC rule requiring disclosure of AI use in MD&A drafting, as SEC Division of Corporation Finance Staff Bulletin No. 20 (2024) addresses AI in business operations disclosures but not in the drafting process itself. That remains an open regulatory question as of September 2026.
Human Review Checklist for AI-Drafted MD&A
Before filing any AI-assisted MD&A, disclosure counsel and the CFO should verify each of the following:
Specificity and forward-looking content
- Does each trend discussion name the specific driver, not just the category? ("Gross margin declined 180 basis points due to higher aluminum input costs in the North America segment" not "margins were impacted by cost pressures.")
- Does the MD&A include management's forward-looking assessment of each material trend, not just a description of what happened?
- Has the opening overview been updated to reflect the current period's most important themes, not carried forward from the prior period?
Trend carry-forward
- Pull the prior period's MD&A. Identify every trend that was disclosed. Confirm each is either updated, continued, or explicitly resolved in the current filing.
- Flag any trend that was present in Q3 but absent from the annual report without explanation.
Earnings call and earnings release consistency
- Read the Q&A transcript from the most recent earnings call. Identify any material issue raised by management or analysts. Confirm it is addressed in the MD&A.
- Cross-reference every quantitative claim in the MD&A against the earnings release. Resolve any discrepancy before filing.
XBRL consistency
- Verify that every percentage or dollar figure stated in narrative form matches the XBRL-tagged financial data exactly, including rounding conventions.
KPI and non-GAAP completeness
- For every KPI disclosed, confirm the three required elements are present: definition and calculation, why useful to investors, how management uses it.
- If any KPI calculation methodology changed from the prior period, confirm the change, reason, and effect are disclosed.
Section placement and PSLRA
- Confirm that forward-looking statements with meaningful cautionary language are in the MD&A, not migrated to footnotes.
- Confirm the contractual obligations section uses the post-2021 principles-based format, not the deprecated tabular format.
Template and vintage check
- Identify the training data vintage or template source for your NLG tool. If it predates February 2021, audit the contractual obligations and liquidity sections against current Item 303 requirements.
The Verdict: Where to Use NLG and Where to Keep Humans in the Lead
NLG tools are genuinely useful for MD&A drafting at the first-draft stage, specifically for the mechanical work: populating period-over-period numerical comparisons, generating the results of operations table narrative, and structuring the liquidity section from cash flow data. These are tasks where internal consistency with financial data is the primary requirement and where NLG performs well.
NLG tools are structurally unsuited, regardless of model capability, for the tasks that Item 303 most demands: forward-looking trend analysis, management's interpretation of causes and implications, trend carry-forward from prior periods, and consistency with earnings call disclosures. These require human judgment and human context that no current AI system has access to.
The SEC's new AI reviewer is calibrated to detect the output of NLG systems. Filing AI-drafted MD&A without a structured human review process does not just risk a comment letter. It risks being surfaced to the Financial Reporting and Accounting Unit for enforcement review. The checklist above is the minimum; for teams that want to benchmark their MD&A specificity against peers before filing, EDGAR full-text search is a practical starting point.
FAQ
What does MD&A stand for in a financial statement? MD&A stands for Management's Discussion and Analysis of Financial Condition and Results of Operations. It is required by Item 303 of Regulation S-K for all public companies filing periodic reports with the SEC.
What does the SEC guidance say companies should do at the beginning of their MD&A? The SEC's 2003 interpretive release (Release No. 33-8350) advocates layered disclosure: the most important themes and highlights at the beginning, with additional detail following. The opening overview should cover material opportunities, challenges, and risks, and must be updated each period. Boilerplate introductions that carry forward unchanged are explicitly flagged as insufficient.
What makes a good MD&A disclosure? A good MD&A lets investors see the company through the eyes of management. It explains the "why" behind the numbers, not just what the numbers show. It quantifies the impact of disclosed factors, carries forward trend disclosures from prior periods, and is consistent with earnings releases and call transcripts. PwC's Viewpoint guidance notes that MD&A should not simply repeat or recalculate information from the financial statements.
Will the SEC's AI flag my MD&A if I used an AI drafting tool? The SEC has not confirmed it screens for AI authorship specifically. What its AI reviewer detects are the patterns AI-drafted MD&A produces: high semantic similarity to peer boilerplate, low specificity, inconsistency between narrative claims and XBRL data, and missing trend carry-forwards. A well-reviewed AI-drafted MD&A that addresses these patterns is not inherently more at risk than a poorly reviewed human-drafted one.
Is there an SEC rule requiring disclosure of AI use in MD&A drafting? No. As of September 2026, no SEC rule or staff guidance specifically requires disclosure of AI use in the MD&A drafting process. SEC Staff Bulletin No. 20 (2024) addresses AI in business operations disclosures, not drafting process disclosures. This remains an open regulatory question that CFOs and general counsel are actively monitoring.
What are the most common SEC comment letter criticisms of MD&A? MD&A is consistently the most-commented section of 10-K filings, appearing in approximately 35 to 45% of all comment letters that include substantive accounting or disclosure comments, per Audit Analytics data. The most frequent staff criticisms are insufficient specificity in trend analysis, failure to quantify the impact of disclosed factors, boilerplate language that does not reflect company-specific circumstances, and inconsistency between MD&A and other sections of the filing. These are precisely the failure modes NLG output is most prone to produce.







