AI Journal Entry Testing and SOX 404 Controls: A 2026 Practitioner Walkthrough
If your team is rolling out AI journal entry testing and calling it a SOX 404 control, your external auditors will have questions you may not be ready to answer. This walkthrough is for controllers, internal audit directors, and SOX program leads who need to design, document, and defend an AI-assisted journal entry control that survives PCAOB scrutiny, not just understand that AI can test 100% of entries.
Key takeaway: The question in 2026 is not whether AI can test journal entries. It demonstrably can. The question is whether your control design, human oversight model, and documentation will satisfy PCAOB AS 2201 and AS 2401 when an inspector looks at it.
For the broader SOX 404 compliance framework, see Finrep's SOX 404 Compliance Checklist and the companion SOX 404 Checklist for AI-Assisted Controls. This article focuses on the practitioner steps: how to build the control, classify it, govern the model, and handle the failure modes.
What PCAOB AS 2401 Actually Requires for Journal Entry Testing
AS 2401 is the regulatory baseline, and it has not been updated for AI. Originally adopted as SAS 99, PCAOB AS 2401 requires auditors to identify and test journal entries and other adjustments for fraud risk. It specifies the characteristics that make an entry high-risk:
- Posted by unusual or unexpected users
- Posted at unusual times (late at night, on weekends, during close)
- Unusual debit/credit combinations
- Entries to accounts that are rarely used or inconsistent with the transaction type
- Round-number entries that suggest estimation or fabrication
AI anomaly-detection models are well-suited to operationalize exactly these criteria at 100% population coverage. The problem is that most implementations do not explicitly map the model's detection logic to these AS 2401 risk factors. If your auditor cannot trace each AI flag category back to a specific AS 2401 criterion, they cannot evaluate whether the control addresses the relevant fraud risks. That is a control design gap, not a technology limitation.
The practical step: Document a crosswalk table in your Risk and Control Matrix (RCM) that maps each AI detection rule or anomaly score dimension to the corresponding AS 2401 risk factor. This single artifact answers the auditor's first question before they ask it.
How to Classify an AI Journal Entry Control in Your Control Matrix
This is the most consequential design decision you will make, and it determines your entire ITGC testing scope.
PwC's 2025 guidance on AI in internal controls identifies three classifications:
| Classification | How it works | ITGC dependency | Audit evidence required |
|---|---|---|---|
| Fully automated control | AI makes and executes the decision | Highest, change mgmt, access, computer ops | System-generated output logs; ITGC testing of the AI system |
| IT-dependent manual control | AI flags; human reviews and decides | Moderate, input data validation; change mgmt | Human sign-off documentation; evidence of exception review |
| Monitoring activity | AI supplements but does not replace a key control | Lower, but still requires data integrity validation | Evidence that exceptions were escalated to a key control |
Most journal entry AI implementations in 2026 are IT-dependent manual controls: the model scores and flags entries, and a human reviewer makes the final disposition decision. Classifying the control correctly matters because:
- A fully automated control requires ITGC testing of the AI system itself under AS 2201, change management, logical access, and computer operations controls all come into scope.
- An IT-dependent manual control still requires ITGC coverage of the data pipeline feeding the model, but the human sign-off is the primary evidence of operating effectiveness.
- A monitoring activity does not replace a key control, if you classify it this way, you still need a separate key control covering journal entry risk.
Get this wrong and you will either over-scope your ITGC testing or, worse, present a monitoring activity as a key control and receive a design deficiency finding.
What "Human in the Loop" Actually Means for Auditors
"Human oversight" is not a philosophy, auditors want specific artifacts. Grant Thornton's 2026 analysis specifies three design requirements for effective human oversight in AI-assisted SOX controls:
- Human-approved control logic with clear criteria for when AI can auto-clear versus route to review
- Transparent explanations behind each AI decision
- Segregation of duties enforced across AI-assisted workflows
In practice, auditors will ask for evidence of all three. Here is what that looks like as documentation artifacts:
For auto-clear thresholds:
- A documented, management-approved policy stating the risk score threshold below which entries are auto-cleared, the rationale for that threshold, and the approval date
- Evidence that the threshold was reviewed and reaffirmed at least annually
For exception review:
- A workflow log showing each flagged entry, the reviewer identity, the review date, and the disposition decision
- For a sample of high-risk flags, evidence that the reviewer actually investigated (not just clicked "approved")
For explainability:
- Model output that shows, for each flagged entry, which risk factors triggered the flag, not just a score
- This is what Grant Thornton means by "transparent explanations behind each AI decision"
The PCAOB's 2024 Annual Report on Inspections flagged recurring deficiencies in auditor testing of automated controls, specifically insufficient evaluation of ITGCs and over-reliance on management representations about system functionality. Your documentation needs to give auditors independent evidence, not assertions.
Model Governance as an ICFR Control: The Four Requirements
The AI model itself is part of your ICFR control environment. Deloitte's 2025 SOX technology survey found that 67% of large accelerated filers are piloting or deploying AI in at least one SOX process, but only 23% have updated their control documentation to reflect AI-specific governance requirements like model validation and drift monitoring. That gap is where deficiency findings will come from.
Map your model governance to COSO 2013 Principle 11 (general control activities over technology) and Principle 16 (ongoing and separate evaluations). The NIST AI RMF Govern-Map-Measure-Manage structure provides a practical overlay that auditors and regulators recognize.
Four governance requirements to document:
1. Model inventory and approval gate Maintain a formal model inventory that records the model name, version, owner, training data vintage, and production approval date. No model enters production without documented sign-off from both the control owner and IT.
2. Change management controls Any change to model logic, thresholds, or training data is a change to the control. Route it through your ITGC change management process, the same process you use for ERP configuration changes. The AICPA's 2023 guidance on AI in accounting explicitly recommends treating AI models as IT applications subject to the same ITGC framework.
3. Drift monitoring and re-validation triggers KPMG's 2025 technical guidance identifies model drift as a specific ICFR risk: a model trained on pre-acquisition or pre-ERP-migration data may degrade materially after a significant business change. Define formal re-validation triggers tied to business events:
- Acquisitions or divestitures above a materiality threshold
- ERP or general ledger system migrations
- Significant accounting policy changes
- Organizational restructurings that change the chart of accounts
Document the re-validation procedure and retain the output as evidence.
4. Training data and input data validation EY's 2025 analysis of PCAOB inspection findings notes that inspectors are increasingly scrutinizing the completeness and accuracy of data used in automated controls. If the data pipeline feeding your AI model has not been validated, the control is considered unreliable regardless of model sophistication. Document the data lineage from the general ledger to the model input, and include completeness and accuracy checks as a control activity.
The Segregation of Duties Problem AI Creates
This is the gap that most SOX teams have not caught yet. When the same AI system both scores entries for risk AND has a configured threshold below which entries are auto-cleared without human review, the system performs both the control execution and the control monitoring function. Under COSO 2013 and AS 2201, this is analogous to a single individual performing incompatible duties.
The specific SOD conflict: the AI model is both the preparer (it processes and classifies entries) and the reviewer (it clears entries below the threshold). No human sees the auto-cleared population.
Grant Thornton's 2026 guidance identifies segregation of duties enforcement across AI-assisted workflows as a design requirement, not an optional feature.
Two mitigations that auditors will accept:
-
Independent monitoring of the auto-cleared population. A separate control owner, someone with no role in configuring the AI model, reviews a random sample of auto-cleared entries each period. Document the sample size, selection methodology, and findings. This is the compensating control.
-
Restrict auto-clearance to a genuinely low-risk population. Set the auto-clear threshold conservatively and document the risk basis. Entries above a defined materiality level, entries to sensitive accounts (related-party, intercompany, top-side), and entries posted outside normal business hours should never be auto-cleared.
Note also that AI cannot substitute for a human in a fraud-deterrence SOD control. As the FloQast analysis points out, SOD as a fraud deterrent requires mutual human accountability, AI is not a party that can be held accountable in the same way.
The Systematic Error Risk: Why SAB 99 Still Applies
This risk does not exist in traditional sampling-based testing, and almost no one is talking about it.
When you sample journal entries manually, errors are random. When an AI model misclassifies entries, errors can be systematic: the model may consistently fail to flag a specific category of transaction, related-party entries structured in a particular way, or intercompany eliminations posted by a specific subsidiary, across every period.
Under SEC Staff Accounting Bulletin No. 99, materiality analysis must consider the aggregate effect of individually immaterial items. A systematic model error that misses the same entry type repeatedly could produce a material aggregate misstatement even if each individual entry is below threshold.
The practical implication: your model validation procedure must include a test for systematic bias, not just overall accuracy. Run the model against a historical population where you know the ground truth and check whether error rates are uniform across entry types, posting users, and account categories. Document the results.
Vendor AI Platforms: The SOC 1 Type II Requirement
If you are using a third-party AI platform for journal entry testing, that vendor's system is part of your ICFR control environment. The AICPA's 2023 guidance recommends treating AI models as IT applications subject to the full ITGC framework, which means the vendor's controls over their system are your controls.
Steps to take:
- Obtain the vendor's SOC 1 Type II report and review it for the control objectives relevant to your journal entry control (data integrity, change management, logical access, availability).
- Identify any complementary user entity controls (CUECs) in the SOC 1 report and confirm you have implemented them.
- If the vendor does not have a SOC 1 Type II report, document your alternative procedures for evaluating their control environment, and be prepared to defend that decision to your auditor.
- Review the SOC 1 report's coverage period against your fiscal year. A report with a June 30 period end does not cover your December 31 year-end without a bridge letter or additional procedures.
For vendor evaluation criteria beyond SOC 1, see Finrep's AI Tools for SOX Compliance: 2026 Evaluation Framework.
The EU AI Act Dimension for US Multinationals
The EU AI Act (effective August 2024, with phased compliance deadlines through 2027) classifies AI systems used in financial services for fraud detection and financial reporting as potentially high-risk. US multinationals with EU operations need to assess whether their AI journal entry system falls under this classification.
If it does, high-risk classification requires:
- Conformity assessments before deployment
- Human oversight mechanisms (which maps directly to the "human in the loop" requirements above)
- Technical documentation of the model
- Registration in the EU AI database
This is not a theoretical concern. An AI journal entry system deployed across a US parent and EU subsidiaries almost certainly processes EU-resident financial data and may fall within scope. Document your assessment either way.
What to Do When the Auditor Disagrees with Your Control Design
This will happen. The PCAOB's December 2024 inspection priorities letter identified AI-assisted controls as a 2025 inspection focus. Auditors are under pressure to scrutinize these controls carefully, and their firm-level guidance on AI control evaluation is still evolving.
If your external auditor proposes a significant deficiency or material weakness related to your AI journal entry control:
-
Ask for the specific AS 2201 basis. Is the proposed deficiency about control design (the control as designed cannot prevent or detect a material misstatement) or operating effectiveness (the control did not operate as designed)? The remediation path is different.
-
Do not concede without analysis. A proposed deficiency is not a final finding. You have the right to present compensating controls and alternative evidence. Document your position in writing.
-
Assess aggregate materiality under SAB 99 and SAB 108. If the auditor's concern is that the AI model missed certain entries, quantify the population and assess whether the aggregate effect is material. A well-documented SAB 99 analysis can support a conclusion that the deficiency does not rise to a material weakness.
-
Check the SEC cybersecurity disclosure interaction. Under the SEC's December 2023 cybersecurity disclosure rule, a failure or breach of an AI ICFR system could trigger both a cybersecurity incident disclosure and a material weakness assessment. If the deficiency finding stems from a system failure rather than a design gap, assess both disclosure obligations simultaneously.
Implementation Checklist: AI Journal Entry Controls for SOX 404
Use this before your next external audit cycle begins.
Control design
- AS 2401 crosswalk: each AI detection rule mapped to a specific AS 2401 high-risk criterion
- Control classification documented: fully automated, IT-dependent manual, or monitoring activity
- Auto-clear threshold documented, management-approved, and annually reaffirmed
- Sensitive account exclusions defined (related-party, intercompany, top-side entries never auto-cleared)
- SOD analysis completed: auto-cleared population covered by an independent compensating control
Model governance
- Model inventory entry created with version, owner, training data vintage, and approval date
- Change management process defined and aligned to existing ITGC change management controls
- Re-validation triggers documented (acquisitions, ERP migrations, policy changes)
- Training data and input data pipeline validated for completeness and accuracy
- Systematic bias test completed and documented
Human oversight evidence
- Exception workflow log maintained: reviewer, date, disposition for each flagged entry
- Sample of high-risk exception reviews documented with investigation evidence
- Model explainability output retained (which risk factors triggered each flag)
Vendor management
- SOC 1 Type II report obtained and reviewed
- CUECs identified and implemented
- SOC 1 coverage period assessed against fiscal year-end
Regulatory
- EU AI Act applicability assessment documented (for multinationals with EU operations)
- SEC cybersecurity disclosure interaction assessed for AI system failure scenarios
- Transition-year documentation: effective date of the new control recorded in the RCM
FAQ
Does AI journal entry testing satisfy PCAOB AS 2401, or do auditors still require traditional sampling? AI testing can satisfy AS 2401 requirements, but only if the model's detection logic is explicitly mapped to AS 2401's specific high-risk criteria (unusual users, timing, accounts, debit/credit combinations, round numbers). Auditors do not automatically accept "we test 100% of entries" as sufficient, they evaluate whether the control addresses the relevant fraud risks.
How do we classify an AI journal entry control in our RCM? Most implementations are IT-dependent manual controls: AI flags, a human decides. Fully automated controls (AI clears without human review) carry the highest ITGC dependency and require the most extensive audit evidence. The classification determines your ITGC scope and the nature of evidence auditors will require.
What does "human in the loop" mean in practice? At minimum: a documented auto-clear threshold with management approval, a workflow log showing reviewer identity and disposition for each flagged entry, and model output that explains which risk factors triggered each flag. A spreadsheet of flagged items with no sign-off evidence is not sufficient.
How do we handle model drift after an acquisition or ERP migration? Define formal re-validation triggers tied to business events and document the re-validation procedure. KPMG's 2025 guidance specifically identifies this as an ICFR risk. A model trained on pre-migration data may degrade materially without anyone noticing.
Does using a third-party AI vendor change our ITGC obligations? Yes. The vendor's system is part of your ICFR control environment. Obtain their SOC 1 Type II report, implement any CUECs, and assess the report's coverage period against your fiscal year-end. If no SOC 1 exists, document your alternative evaluation procedures.
What if the external auditor proposes a deficiency finding on our AI control? Ask for the specific AS 2201 basis (design vs. operating effectiveness), present compensating controls and evidence, and run a SAB 99 materiality analysis on the population the auditor believes was not covered. Also assess whether the SEC cybersecurity disclosure rule is triggered if the issue stems from a system failure.







