Abstract AI legal delivery architecture transforming source preparation and expert review into an accepted output
|

Cheap Tokens Do Not Make Legal Work Cheap

A model can draft a contract clause for cents while the firm still loses money delivering it.

That is the accounting error hiding inside many legal AI programs. The model invoice is visible, so it becomes the headline. The expensive work sits around it: preparing client material, checking authority, resolving ambiguity, reviewing risk, fixing exceptions, explaining the result and absorbing write-offs.

A fast first draft is not a profitable matter. It is only one event inside a production system.

Here’s what works: measure the economics at the accepted deliverable, not the model response. Build a Matter Unit-Economics Sheet for one repeatable matter type and run it for 30 days. If cycle time, realization and margin do not improve after full review and rework are counted, redesign the workflow or stop it.

I have worked with automation since 2003 and spent more than 20 years building hosting and software infrastructure, including scaling a business from €600,000 to €240 million ARR. Cost has a habit of leaving one queue and reappearing in another. Production economics become real only when every handoff, exception and recovery path is measured.

Cheap inference creates an expensive illusion

AI vendors price tokens, seats and API calls because those are easy to meter. A law firm sells something different: a defensible output delivered under professional responsibility, client instructions, commercial terms and a deadline.

The gap between those units is where margin disappears.

A contract-review workflow may show a model cost of €0.40 per document. That number says nothing about the 18 minutes spent cleaning the source file, the partner who checks a non-standard indemnity, the associate who reruns retrieval after a missing schedule, or the time written off because the client expected the work to be cheaper.

The same pattern appears in tax, accounting and consulting. The machine step gets faster. The surrounding system remains fragmented. Work moves from drafting into preparation, verification and exception handling, then vanishes inside unstructured time entries.

The evidence says professional services has moved beyond experimentation. Thomson Reuters’ 2026 AI in Professional Services report says organizational generative-AI use rose from 22% to 40%. Yet only 18% of professionals said their organizations track ROI, while another 40% did not know whether ROI was measured.[1]

Thomson Reuters’ Future of Professionals Report 2026 frames the next divide clearly: AI is embedded in professional work, but value depends on execution and the right infrastructure. The gap is widening between organizations producing impact and those falling behind across clients, talent and governance.[2]

Client pressure adds a second constraint. Litera’s 2026 State of Legal AI research reported that 85% of surveyed firms felt or expected direct client pressure around AI strategy, while 32% could not confidently demonstrate AI value to an important client.[3] That is vendor-sponsored survey evidence, not a census of every firm. But it exposes the operating problem: adoption is moving faster than the ability to prove value and explain commercial treatment.

You cannot solve that with another usage dashboard.

The Matter Unit-Economics Sheet

The Matter Unit-Economics Sheet is a production record for one accepted output. It connects the full cost of delivery to the fee actually realized.

Use one row per deliverable and capture twelve fields:

  1. Matter type. Define a narrow, repeatable unit such as first-pass NDA review, lease abstraction, due-diligence issue list or standard research memo.
  2. Source-preparation minutes. Count document cleaning, OCR, deduplication, metadata repair, client follow-up and knowledge-base preparation.
  3. Model and tool cost. Include inference, retrieval, document processing, licensed data and workflow-platform usage.
  4. AI execution time. Measure elapsed time from accepted input to usable first output, including failed runs and retries.
  5. Expert-review minutes. Record the lawyer, accountant or consultant time required to verify the work—not the optimistic review allowance in the business case.
  6. Exception type. Classify missing sources, conflicting instructions, unusual clauses, jurisdiction issues, retrieval gaps, formatting failures and policy blocks.
  7. Rework minutes. Separate avoidable production repair from required professional judgment.
  8. Cycle time. Measure the complete path from ready input to client-accepted output.
  9. Quality result. Use a matter-specific acceptance test: missed issues, citation accuracy, correction rate, escalation rate or client rejection.
  10. Realized fee. Capture what the firm actually bills and collects, not the standard rate multiplied by recorded hours.
  11. Write-off or scope leakage. Record discounts, unbilled review, extra variants and work absorbed under a fixed fee.
  12. Contribution margin per accepted deliverable. Subtract direct professional time, model/tool cost and allocated delivery overhead from realized revenue.

The phrase accepted deliverable matters. A generated draft that fails review is work in progress, not output. A memo that requires two senior rewrites is not a successful automation because the first version arrived quickly. A contract review that saves associate time but adds partner time may still be valuable—but only if the economics and risk profile support the trade.

This sheet forces the firm to see the entire path.

Matter Unit-Economics Sheet framework

Measure the baseline before the AI version

Do not compare the new workflow with a story about how work used to happen. Sample actual completed matters.

Choose one matter type with enough repetition to produce a useful comparison. Twenty completed deliverables is usually sufficient for a 30-day operational proof—not to claim universal statistical certainty, but to expose where time, quality and margin move.

For each baseline matter, reconstruct preparation time, drafting time, review time, exceptions, cycle time, realized fee and write-off. Time-entry data will be incomplete. That is useful information. Interview the people who performed the work and mark every estimate as an estimate rather than polishing it into false precision.

Then define the acceptance condition. For an NDA review, it might be: all mandatory clauses checked against the approved playbook, every deviation cited to the source text, no unresolved high-risk item, and named lawyer approval before client delivery.

Without that gate, teams optimize for first-draft speed because it is the easiest metric to improve. The business needs accepted-output economics.

Separate judgment from rework

Professional judgment is not a defect to automate away.

A partner deciding whether a liability cap is commercially tolerable is doing the work the client hired the firm to perform. An associate manually finding a clause the retrieval system missed is repairing the production system. Both consume time, but they mean different things.

Tag review minutes in two buckets:

  • Required judgment: interpretation, negotiation posture, risk acceptance, client context and final professional approval.
  • Avoidable rework: missing sources, poor extraction, unsupported statements, repeated formatting, duplicated review, prompt repair and workflow failures.

The first bucket protects quality and often carries the highest client value. The second is an engineering backlog.

This distinction also prevents a damaging management response. If all human review is treated as inefficiency, teams will hide necessary judgment or push unsafe work downstream. If all review is treated as professional necessity, nobody fixes a weak workflow.

Here’s the operator rule: preserve judgment; attack rework.

Calculate cost at median and P95

Averages make unstable workflows look healthy.

Suppose 16 reviews need ten minutes, three need 35 minutes and one needs 110 minutes. The average may still look attractive. The exception at the tail is where deadlines, write-offs and partner frustration live.

Track both median and P95 for preparation, expert review, rework and cycle time. Then show the exception rate by type.

The unit economics can stay simple:

Contribution margin per accepted deliverable = realized fee − professional labor − model/tool cost − allocated delivery overhead.

Use loaded labor cost for internal economics, not the billing rate. Keep revenue treatment separate. A fixed-fee matter, capped arrangement and hourly engagement react differently to saved time. Faster delivery can improve margin under a fixed fee and reduce revenue under hourly billing unless capacity is redeployed or pricing changes.

Allocate delivery overhead consistently as well. Include workflow maintenance, knowledge-base curation, security review, monitoring and the support capacity required when a run fails near a deadline. Do not load the entire innovation budget onto one matter type, but do not pretend the production service operates for free. A simple monthly allocation based on accepted volume is more honest than excluding infrastructure because another department pays the invoice.

That is why “hours saved” is not an ROI model. It is an operating input.

The 30-day proof path

Days 1–5: choose and baseline

Select one bounded matter type. Pull 20 recent accepted deliverables. Reconstruct the full production path and agree the acceptance test with the responsible partner.

Assign one workflow owner and one economic owner. Innovation can run the system; finance or practice operations must validate the cost logic.

Days 6–10: instrument the path

Add event capture at every handoff: source ready, AI started, first output, review started, exception opened, rework completed, partner approved and client accepted.

Do not ask professionals to write essays. Use timestamps, short exception codes and a required reviewer field. The system should create the row automatically wherever possible.

Days 11–20: run in shadow, then live

Process the first five deliverables in shadow mode against the existing method. Compare missed issues, unsupported statements, preparation time and review burden.

If the acceptance gate holds, move the next ten to live production with named human approval. Preserve the old path as a recovery option. Build-Operate-Transfer works because the firm owns the process definition, evidence and rollback—not because it rents a clever prompt.

Days 21–25: price the exceptions

Rank exception types by frequency, minutes and commercial impact. Fix the top engineering failure rather than rewriting the entire workflow.

If missing schedules create most rework, repair intake. If citations fail, fix retrieval and source boundaries. If partner review varies wildly, define a clearer playbook before blaming the model.

Days 26–30: make the decision

Compare baseline and live results for median and P95 cycle time, quality, realization and contribution margin.

Then choose one outcome:

  • Scale when quality holds and accepted-deliverable margin improves.
  • Redesign when the economic gain is plausible but one or two exception classes dominate.
  • Stop when review, rework or commercial leakage consumes the apparent saving.

Stopping is not failure. It is proof that prevented a weak workflow from spreading across the firm.

What managing partners should ask on Monday

Do not ask how many people used the AI tool last week. Ask five sharper questions:

  1. Which matter type has an agreed accepted-output definition?
  2. What is the full cost per accepted deliverable at median and P95?
  3. How much review is required judgment versus avoidable rework?
  4. Did the realized fee, write-off and contribution margin improve?
  5. Who owns the next exception reduction—and when will it ship?

Those questions change the internal conversation. AI stops being a technology initiative and becomes a delivery system with economics, controls and accountable owners.

Cheap tokens are useful. They lower one component cost and make more workflows worth testing. But they do not make legal work cheap by themselves. The firm captures value only when the complete production path becomes faster, more predictable and more profitable without weakening professional judgment.

That is the standard: 30 days to proof, measured at the accepted deliverable.

Book a 30-minute strategy call

Sources

[1] Thomson Reuters — 2026 AI in Professional Services Report
[2] Thomson Reuters Institute — Future of Professionals Report 2026
[3] Litera — 2026 State of Legal AI

Similar Posts