CPA · CISA · CISM · CDPSE · CCSE · MBA
Continuous controls monitoring dashboard showing AI-assisted anomaly detection across transaction data

AI and Agentic Automation for Audit and Continuous Monitoring

Full-population testing, anomaly detection and continuous controls monitoring, built so an auditor will actually rely on the output.

What can AI audit automation actually do today?

Test full populations instead of samples, detect anomalies across transaction sets, and monitor controls continuously. What it cannot do is audit a company. The output has to be built so an auditor will rely on it, which means documented logic, reproducible results and evidence a reviewer can follow.

The vendor pitch oversells this and the sceptical response dismisses it. The genuinely useful territory is narrower than the first and wider than the second.

AI audit automation is currently oversold in one direction and dismissed in the other. The vendor pitch implies an agent that audits your company. The sceptical response is that none of it is admissible as evidence. Both are wrong, and the useful territory in between is narrower and more specific than either.

What genuinely works today: testing full populations instead of samples, detecting anomalies across transaction volumes no team could review manually, monitoring controls continuously rather than annually, and drafting documentation that a human then corrects. What does not work yet: an AI conclusion standing as audit evidence without a controlled, reproducible process behind it.

AI audit automation in practice: from sampling to full population

Traditional audit testing samples because examining everything was impractical. That constraint has largely disappeared for structured data, and its disappearance changes what testing can conclude.

A sample of forty journal entries supports a statistical inference about the population. Testing every journal entry posted in the year, filtered for the characteristics that indicate risk: entries posted by unexpected users, entries to unusual account combinations, round-dollar amounts, entries posted outside business hours, entries at period end reversing shortly after, identifies specific items rather than inferring a rate.

This is not new in principle; data analytics in audit predates the current wave by decades. What language models change is the cost of building the analysis. Work that required a specialist and several weeks can now be specified, built and refined in days, which moves it from something reserved for large engagements to something available on ordinary ones.

Where language models genuinely add capability

The specific advantage is unstructured data, the material that resisted automation entirely.

  • Contract review at scale. Extracting revenue recognition terms, change-of-control provisions, termination rights, liability caps and data-processing clauses across hundreds of agreements, producing a structured dataset from documents that were previously read one at a time or not at all.
  • Free-text field analysis. Journal entry descriptions, expense narratives, support tickets and vendor names contain signal that keyword search misses because it depends on meaning rather than wording.
  • Policy-to-practice comparison. Reading a written procedure and comparing it against evidence of what was actually done, at a volume that makes systematic comparison feasible.
  • First-draft documentation. Process narratives, control descriptions and system descriptions drafted from interview transcripts and system evidence, then corrected by a person. This is straightforwardly a time saving, and it is where most firms should start.

Continuous controls monitoring

Annual controls testing tells you the control worked on the days you sampled. Continuous monitoring tells you whether it is working now, which is a different and more useful proposition, particularly for SOX programmes and SOC 2 observation windows, where a control failing in month two and discovered in month eleven is an exception in the report.

Controls that automate well: segregation of duties conflicts detected as they arise rather than in a quarterly review, new vendor creation matched against the payment file and against employee address data, access changes reconciled to approval records, journal entries flagged against defined risk criteria, and expense claims tested against policy thresholds.

The design constraint that determines whether this succeeds is alert volume. A monitoring system generating two hundred alerts a week gets ignored within a month, and an ignored control is worse than no control because it creates documented false assurance. Tuning to a volume the responsible person can genuinely review (and defining what happens to each alert) is most of the implementation work.

Making AI-assisted work defensible

The question every auditor and regulator will ask is how you know the output is right. The answer cannot be that the model is usually accurate.

What makes AI-assisted procedures defensible is the same thing that makes any procedure defensible, a controlled, documented, reproducible process:

  • Deterministic where it matters. The extraction, filtering and matching logic is code, not a prompt. Language models are used for classification and extraction; conclusions come from rules that produce the same output every time.
  • Validated against a known-answer set. Before relying on an automated procedure, it is run against a population where the answer was established manually, and the error rate in both directions is measured and documented.
  • Human review of every exception. The system identifies; a person concludes. That boundary is not negotiable and should be visible in the workpapers.
  • Version and change control over the automation itself. Which version of the logic produced which result, who changed it, when, and who approved it, the same change management discipline applied to any system relied upon.
  • A retained audit trail. Inputs, logic version, outputs and dispositions preserved so the procedure can be re-performed.
  • Data governance. What was sent to which model, under what terms, and whether the provider's agreement permits it. Sending client financial data to a consumer AI service is a confidentiality breach regardless of the quality of the analysis.

What is deliberately not claimed

An AI agent does not form an audit opinion, and any product suggesting otherwise is misrepresenting both the technology and the professional standards. Models produce confident output on questions they have handled poorly, which is precisely the failure mode that matters in assurance work. And a procedure that cannot be explained to an auditor in terms they can evaluate will not be relied upon, however well it performs.

The realistic position: AI reduces the cost of finding things by a large factor, and changes the conclusion process not at all.

Javed Peeran CPA

Javed Peeran

CPA · CISA · CISM · CDPSE · CCSE · MBA

Licensed by the California Board of Accountancy and the author of every article published here. Thirty years of practice covering external audit of banks, insurers and mortgage companies, fifteen years as CFO and Corporate Controller inside technology companies, and IT governance and security compliance work spanning SOX 404, SOC 1 and SOC 2, ISO 27001, FISMA, FedRAMP, PCI DSS, HIPAA/HITECH, CCPA and GDPR, plus Oracle ERP migrations and, more recently, generative-AI audit automation.

What the engagement delivers

  • Full-population journal entry testing with risk-based filtering criteria
  • Anomaly detection across transaction sets with tuned, reviewable alert volumes
  • Continuous controls monitoring design and implementation
  • Contract and unstructured document analysis at scale
  • Automated segregation of duties conflict detection
  • Vendor master and payment file matching routines
  • Validation testing of automated procedures against known-answer populations
  • Change control and audit trail design for the automation itself
  • AI data governance policy covering permitted tools, data classes and provider terms
  • Workpaper documentation written to withstand external auditor review

How a typical engagement runs

  1. Pick the right first case

    Start where the data is structured, the rule is clear and the manual effort is high, journal entry testing, duplicate payment detection, access reconciliation. Not the hardest problem in the business.

  2. Build deterministic first

    Rules-based logic handles the conclusion; models handle extraction and classification. Reversing that is how projects produce output nobody can defend.

  3. Validate against known answers

    Run against a population already tested manually and measure the error rate in both directions. Without this step there is no basis for reliance.

  4. Control and hand over

    Version control, change approval, retained audit trail, and training so the client’s team operates it. Automation that requires the consultant to return annually was built wrong.

AI & Agentic Audit Automation across Ventura County and Los Angeles

This service is delivered on site and remotely across the firm's service area. See how it applies locally:

AI & Agentic Audit Automation: questions we are asked

Not answered here? Ask Javed directly

Will an external auditor accept AI-assisted testing?

They will accept a controlled, documented, reproducible procedure, which an AI-assisted procedure can be. What they will not accept is output from a tool nobody can explain.

The practical requirements are consistent: show the logic, show it was validated against a known-answer population, show that exceptions were reviewed by a person, show change control over the automation, and retain the trail so it can be re-performed. Procedures built to that standard are relied upon routinely. Procedures built as a prompt and a spreadsheet are not.

Is it safe to put our financial data into an AI model?

It depends entirely on which model and under what agreement. Enterprise offerings from the major providers generally contractually exclude customer data from training and provide defined retention terms; consumer tiers frequently do not.

This is a governance question before it is a technical one, and it is one of the first things established in an engagement: which tools are permitted, for which data classifications, under which contractual terms. Sending client financial records to an unapproved consumer service is a confidentiality breach irrespective of how good the analysis is.

Where should a mid-market finance team start?

Duplicate payment detection or journal entry analysis. Both use structured data you already have, both have clear rules, both replace work someone currently does manually or does not do at all, and both produce a result that is immediately verifiable, which builds the credibility needed for anything more ambitious.

Where not to start: a general-purpose assistant with access to everything. Broad scope produces impressive demonstrations and no measurable outcome.

Does this replace our finance team?

No, and the engagements that assume it will tend to fail. What it replaces is a specific category of work, reconciling large files, scanning transaction listings for exceptions, reading documents to extract a handful of fields.

The conclusion, the judgment about materiality, the conversation with the person who posted the unusual entry, and the decision about whether an exception matters all remain with people. In practice teams end up doing more analysis rather than less, because the analysis became affordable.

How do you keep alert volumes manageable?

By tuning against observed behaviour before enabling anything, and by setting the threshold at what the responsible person can genuinely review each week rather than at what the rule technically catches.

The failure pattern is well-established: a monitoring system launched with untuned thresholds generates a flood, the reviewer stops reading within weeks, and the organisation now has a documented control providing false assurance. Tuning is not a refinement phase; it is the implementation.

Related services

SOX 404 Compliance

Scoping, control design, testing and remediation for Section 404, including the first-year programmes that decide whether year two is manageable.

SOX 404 support

Organisations we have worked with

Three decades of audit, controls and finance leadership across banking, card, mortgage, insurance, staffing and semiconductor.

  • Diodes Incorporated
  • City National Bank
  • Robert Half
  • SMBC
  • PennyMac
  • American Express
  • Zenith Insurance
  • Capco Consulting Services
  • WebVision

Get in touch

Enquire about ai & agentic audit automation

A sentence or two about your situation (the standard involved, the deadline, and what has already been attempted) is enough to get a useful reply.

Have a deadline, or just a question?

Send the shape of it. The first call is diagnostic, not billed, and it regularly ends with a smaller engagement than the one you asked about.

Javed Peeran CPA Request a consultation

Answered personally, within one business day. Your details are used only to reply to you, see our privacy policy.

Talk through a ai & agentic audit automation engagement

Thirty years of audit, financial leadership and IT governance in one engagement, and a direct answer about scope, sequence and cost before anything is signed.

WhatsApp Us
Call Now