Abstract AI diligence architecture revealing exception queues and dependency risk beneath a controlled path to validated enterprise value
|

In AI Diligence, Ask Who Empties the Exception Queue

In AI diligence, ask who empties the exception queue

An “AI-enabled” target can look excellent in a management presentation. Revenue per employee is up. Headcount grew slower than sales. Customer response times fell. A dozen workflows are described as automated.

Then two operations managers leave after close, the model vendor changes an API, and the EBITDA case starts leaking.

The problem usually wasn’t the AI. The buyer underwrote an outcome without underwriting the operating mechanism behind it.

I’ve executed more than 15 acquisitions. One lesson keeps repeating: a result is only as durable as the system, people and controls that produce it. AI makes that diligence question harder because the visible workflow can hide a lot of human labor. A small team may be clearing edge cases, repairing outputs, moving data between systems and quietly preventing customer-facing failures.

That work sits in the exception queue. If nobody measures it, the buyer can mistake deferred operating debt for structural margin.

Here’s what works: treat every material AI-enabled workflow as a small production system. Sample actual runs. Measure straight-through work. Price the exceptions. Map the dependencies. Put the largest remediation item into the 100-day plan.

That is 30 days to proof, not six months to recommendations.

AI adoption is no longer the useful diligence question

AI usage has moved quickly enough that “does the company use AI?” tells you very little. The Stanford AI Index 2025 reported that 78% of organizations used AI in 2024, up from 55% a year earlier. McKinsey’s 2025 State of AI also found 78% using AI in at least one business function, with 71% regularly using generative AI in at least one function.

Adoption is becoming table stakes. Repeatable economics are not.

For a PE buyer, the useful questions are operational:

  • Which workflows affect revenue, gross margin, service quality or compliance?
  • What percentage of runs complete without human repair?
  • Which exceptions require expert judgment, and which expose bad process design?
  • Who owns the queue, and what happens if that person leaves?
  • Which model, vendor, prompt, dataset and integration does the result depend on?
  • Can management prove the control with logs, not a slide?
  • What would it cost to make the workflow transferable after close?

This distinction matters because human review is not automatically waste. In legal, clinical, financial or high-value commercial work, expert review may be the product. The diligence task is not to eliminate people. It is to identify where human judgment creates value, where humans compensate for brittle automation, and whether the cost is already reflected in the plan.

The hidden EBITDA bridge

Imagine a target claims its AI quotation workflow saves eight sales-support roles. The headline benefit is easy to model.

Now sample 200 quotations from the last quarter. You find that 62% completed straight through. The remaining 38% required an average of 17 minutes of human intervention. Two senior employees handled most of the complex cases. Their time was booked across sales operations, product and customer success, so the workflow P&L never saw the full cost.

The saving still exists. It is simply smaller, more concentrated and less transferable than the headline suggests.

That changes the investment case in four places:

  1. Quality of earnings: hidden exception labor belongs in the normalized cost base.
  2. Key-person risk: knowledge may sit with the people who know which AI outputs not to trust.
  3. Capex and 100-day planning: remediation needs an owner, budget and sequence.
  4. Exit readiness: the next buyer will ask whether the margin improvement is controlled and repeatable.

The hidden door is to build an AI operating-debt schedule alongside the normal technology and quality-of-earnings work. It turns “AI-enabled” from a management adjective into a set of testable operating claims.

The AI Operating-Debt Schedule

Use one row per material AI-enabled workflow. Start with workflows tied to revenue recognition, pricing, customer commitments, service delivery, regulated decisions, financial reporting or meaningful labor savings.

Score each row across eleven fields:

  1. Claimed workflow: What does management say is automated?
  2. Business consequence: Which revenue, cost, customer or control outcome depends on it?
  3. True straight-through rate: What share completes without human correction, rerouting or rework?
  4. Exception load: How many cases need intervention, and how many minutes do they consume?
  5. Judgment class: Is intervention valuable expert judgment or repair work caused by weak design?
  6. Key-person dependency: Who knows how to resolve failures, and is that knowledge documented?
  7. Model and vendor dependency: What breaks if pricing, rate limits, model behavior or API terms change?
  8. Data rights and lineage: Can the company show where inputs came from and whether it can use them?
  9. Control evidence: Are approvals, overrides, failures and customer-visible actions logged?
  10. Monthly run cost: Include models, software, infrastructure, monitoring and human exception time.
  11. Remediation and ownership: What must change, what will it cost, and who owns it in the 100-day plan?

AI Operating-Debt Schedule from management claim through exceptions, dependencies, controls and normalized EBITDA

The framework is deliberately operational. It does not try to produce one magical “AI maturity” score. A target can be advanced in one workflow and dangerously informal in another. Diligence needs to preserve that resolution.

Measure actual runs, not configured capabilities

A demo proves that a workflow can work. Diligence needs evidence that it does work under normal volume, messy data and real exceptions.

Choose a sample period that captures peaks, month-end activity and at least one abnormal event. Pull system logs where possible. Reconcile them to the business system of record. Then classify each run:

  • completed straight through;
  • completed after approved human review;
  • completed after unplanned repair;
  • failed safely with no external consequence;
  • failed with a customer, financial or control consequence;
  • abandoned or completed outside the system.

This gives you a factual denominator. Without it, “90% automated” may mean 90% of steps, 90% of happy-path demos or 90% of transactions before an undocumented manual check.

NIST’s Generative AI Profile treats measurement, monitoring, documented oversight and incident handling as operating practices, not policy decoration. The same logic belongs in investment diligence. If management cannot reproduce the evidence, the control is not yet transferable.

European buyers also have a regulatory reason to care about the operating trail. The official EU AI Act places concrete duties on deployers of high-risk AI systems, including competent human oversight, monitoring and retention of automatically generated logs when those logs are under the deployer’s control. Not every target workflow will be high-risk under the Act. But a diligence process that cannot identify deployed systems, owners, logs and oversight will struggle to classify the exposure correctly.

Normalize exceptions without destroying expert work

The blunt approach is to label every human touch as automation failure. That produces bad diligence and worse post-close decisions.

Split exceptions into three classes:

Class A: valuable judgment

A trained professional reviews a high-impact recommendation, resolves ambiguity or accepts responsibility for a decision. Preserve this. Improve the evidence and capacity model around it.

Class B: operational variance

The workflow meets a legitimate edge case: a non-standard contract, unusual customer configuration or missing external input. Decide whether the case should remain exceptional or become a supported path.

Class C: design debt

Humans repeatedly repair malformed outputs, copy data between disconnected systems, recover from brittle prompts or catch errors that basic validation should block. This is operating debt. Cost it and remediate it.

The distinction protects both margin and quality. “Human in the loop” is not a control if the human is overloaded, untrained or rubber-stamping. Equally, removing expert review to improve a straight-through-rate metric can create a larger commercial or regulatory risk.

Here’s the test: if volume doubled next month, which exception class would break first? That answer is often more valuable than the current automation percentage.

Price dependencies as part of the deal

An AI workflow rarely stands alone. It depends on identity, source data, system permissions, model access, orchestration code, monitoring and a destination system that records the result.

Map those dependencies for the top five workflows. Then run three practical scenarios:

Vendor shock: The primary model becomes unavailable or costs three times more. Can the workflow switch models, degrade safely or stop without corrupting state?

People shock: The two heaviest exception handlers leave. Can another operator resolve cases from documented evidence and runbooks?

Volume shock: Transactions double during a peak month. Do exception minutes scale linearly, or does the queue become a customer-facing backlog?

This is where my 20+ years in hosting and infrastructure shape the view. Customers do not pay for elegant architecture diagrams. They pay for reliable outcomes when components fail. AI infrastructure is no different. Portability, observability and clear ownership create enterprise value because they keep the operating result intact under pressure.

A 30-day proof path for the deal team

You do not need a six-month AI audit. In 30 days, the deal team and operating partner can expose the material operating debt and decide what belongs in valuation, documentation and the 100-day plan.

Days 1–5: select the material workflows

List every AI-enabled workflow management claims affects growth, labor, customer experience or control. Rank by EBITDA impact and downside consequence. Select the top three to five. Name one management owner for each.

Days 6–10: establish the run denominator

Export 30 to 200 recent runs per workflow, depending on volume. Reconcile outcomes to CRM, ERP, ticketing, billing or another authoritative system. Record straight-through completion, planned review, repair, failure and off-system completion.

Days 11–15: time and classify the exceptions

Observe the people handling exceptions. Measure minutes, not guesses. Separate valuable judgment, legitimate variance and design debt. Document the knowledge used to resolve each recurring class.

Days 16–20: map control and dependency debt

Capture model and vendor dependencies, data rights, identities, approvals, logs, monitoring and fallback behavior. Run one vendor-shock tabletop and one key-person-shock tabletop. Record where evidence is missing.

Days 21–25: normalize the economics

Calculate monthly model, software, infrastructure and human exception cost. Adjust the claimed saving for recurring repair work. Estimate one-time remediation cost and the capacity required at forecast volume.

Days 26–30: make the investment decision explicit

Put every material issue into one of four buckets: valuation adjustment, deal protection, 100-day remediation or accepted operating risk. Assign an owner and date. Scale only the workflows that show stable economics and controlled exceptions.

The stop rule is simple: if the target cannot produce a trustworthy run sample, named owner and exception trail for a material claim, do not underwrite the full benefit yet.

The scale rule is equally simple: if two consecutive samples show stable straight-through performance, known exception economics and recoverable dependencies, fund the next deployment.

Underwrite the mechanism, not the label

AI can create real margin, faster decisions and better customer outcomes. It can also hide labor, concentrate knowledge and add dependencies that only become visible after close.

A buyer does not need to be anti-AI to challenge the claim. The opposite is true. Serious AI value creation starts by measuring the production system honestly.

Ask who empties the exception queue. Ask what they know. Ask what happens when they leave. Then put the answer into normalized EBITDA and the 100-day plan.

That is how AI becomes durable enterprise value instead of a footnote in the investment committee deck.

Book a 30-minute strategy call

Similar Posts