If AI Builds the Deck, Measure Decision Density—Not Slides per Day
Claude can produce 60 polished slides before lunch. That does not mean a consulting team produced a useful deliverable before lunch.
It may mean the team created 60 surfaces for a manager to check, 60 chances to repeat the same point, and 60 reasons for a client to ask, “What exactly do you need us to decide?”
Slides per day was never a good proxy for value. Generative AI makes it actively misleading. When the marginal cost of another page approaches zero, page count measures tool output, not professional progress.
A client deck earns its fee when it makes a consequential decision easier to take: the question is explicit, the evidence is traceable, the inference is visible, the options are real, the recommendation is owned, uncertainty is stated, and the next action has a name beside it.
That is decision density.
AI has accelerated production faster than firms measure value
The adoption curve is no longer the interesting part. The 2026 Thomson Reuters AI in Professional Services Report, drawing on more than 1,500 professionals across legal, tax, accounting, risk, fraud and government, reports that organizational GenAI use rose from 22% to 40%. More than 80% of current users engage with it weekly. Yet only 18% say their organizations track ROI, while another 40% do not know whether ROI is measured at all.
Those are vendor-reported survey findings. Thomson Reuters sells AI-enabled professional products, so the report is useful market evidence, not independent proof that a particular tool or operating model creates value. The useful tension is still clear: production is spreading faster than measurement discipline.
An earlier randomized field experiment offers a second signal. In the “Navigating the Jagged Technological Frontier” working paper, researchers assigned 758 Boston Consulting Group consultants to work with or without GPT-4 on realistic consulting tasks. On tasks inside the model’s capability frontier, AI users completed 12.2% more tasks, worked 25.1% faster and produced outputs rated more than 40% higher in quality. On a task deliberately placed outside that frontier, consultants using AI were less likely to reach the correct answer. Co-author Ethan Mollick provides an accessible summary of the experiment and its limitations.
The study was a preregistered working paper produced with BCG, not a peer-reviewed trial of AI-generated client decks. It used GPT-4-era tasks and does not prove that adding AI to a current presentation workflow improves client decisions. It does show why raw speed is an incomplete metric: the same system can improve output inside its capability boundary and confidently pull professionals toward a wrong answer outside it.
More pages can therefore arrive faster, look better and still make the senior-review problem worse.
The hidden cost of the 60-slide draft
Consider a strategy team preparing a steering-committee deck on whether a client should consolidate three service operations.
The old workflow produced 35 slides in five days. The AI-assisted workflow produces 60 in one day. The team reports an 80% reduction in drafting time.
But then:
- a manager spends six hours finding duplicated claims;
- a subject-matter expert discovers that two charts use different cost baselines;
- the partner rewrites the recommendation because the draft never separated evidence from inference;
- the client meeting ends with a request for “a shorter version with the actual decision”; and
- no one records who must approve the consolidation or by when.
The production step improved. The delivery system did not.
The correct comparison is not 35 slides versus 60. It is the review effort and decision movement created by each version.
Volume model
60 generated slides → senior review → compression → meeting → unclear next step
Decision-density model
Client question → verified evidence → explicit inference → options
→ recommendation + uncertainty → decision requested → named owner
The second chain may require eight slides, twelve slides or a one-page memo with an appendix. The format is secondary. The chain is the product.
The PromptPartner Decision-Density Score
The PromptPartner Decision-Density Score (DDS) is a review rubric for each material decision in a deliverable. It is not a scientific index, a billing formula or a claim that one point predicts a percentage of client value. Its purpose is to force a team to inspect the parts that polished prose can conceal.
Score each dimension 0, 1 or 2:
| Dimension | 0 — absent | 1 — present but weak | 2 — explicit and reviewable |
|---|---|---|---|
| Client question | Topic only | Question implied | One decision-relevant question stated in client language |
| Verified evidence | Unsupported claim | Evidence shown without provenance or common baseline | Source, date, scope and common baseline can be checked |
| Explicit inference | Fact and interpretation blended | Reasoning implied | The step from evidence to conclusion is stated |
| Options | Single predetermined answer | Alternatives named but not compared | Viable options use consistent criteria and trade-offs |
| Recommendation | No position | Direction without rationale | Recommended option, rationale and consequence are clear |
| Uncertainty | Confidence performed | Caveat buried | Assumption, range or disconfirming evidence is visible |
| Decision requested | “For discussion” | General agreement requested | Named decision, decision-maker and deadline are stated |
| Action owner | No follow-through | Team or function named | One accountable owner, first action and due date are named |
A complete chain scores 16. Do not average away a zero in verified evidence, decision requested or action owner. Those are stop signs, even if the total looks respectable.
Use three plain review states rather than decimals:
- Draft: one or more mandatory links are absent.
- Reviewable: every link is present, but at least one depends on an unverified source, hidden assumption or vague owner.
- Decision-ready: all eight links are explicit and a senior reviewer can trace the recommendation without reconstructing the analysis.
Then attach two load measures:
- Client-facing pages per material decision. Count only pages needed to understand or take the decision; put supporting detail in a linked appendix.
- Senior review minutes per accepted decision. Measure the scarce professional time required before the client accepts, rejects or advances the recommendation.
There is no universal “correct” page ratio. A regulatory opinion and a pricing recommendation carry different evidence burdens. Establish a baseline for one recurring deliverable type and compare like with like. The objective is not to force every answer onto one slide. It is to remove pages that do not strengthen a decision chain.
A before-and-after review
The 60-slide operations deck contains three apparent recommendations. Let us inspect one: “Consolidate service operations into one regional hub.”
Before
- Fifteen slides describe the current organization.
- Four charts show cost, but one includes transition expense and three do not.
- The recommendation appears in an executive summary.
- Country-level service risk sits in a footnote.
- The final page says “align on next steps.”
It looks finished. Under the DDS rubric, the baseline is missing a common evidence definition, an explicit inference, comparable options, visible uncertainty, a precise decision and an accountable owner. More design work will not fix those gaps.
After
- Question: Should the client consolidate the three operations by Q2, phase them over 18 months, or retain the current model?
- Evidence: A six-month cost and service baseline, using the same allocation rules across all three sites, with source-system links.
- Inference: Consolidation creates the largest run-rate saving only if attrition stays below the stated range and two key controls can move without service interruption.
- Options: Full consolidation, phased consolidation and no structural change, compared on cost, service risk, control readiness and reversibility.
- Recommendation: Phase two operations first; hold the regulated workflow until the control test passes.
- Uncertainty: Attrition and migration effort remain ranges; a named test would invalidate the recommendation.
- Decision requested: The steering committee must approve the first phase and its budget by Friday.
- Owner: The COO owns approval; the regional operations director starts the control test on Monday.
The core deck now uses ten pages plus an evidence appendix. The team has not “won” because it deleted 50 slides. It has won if review time falls, evidence challenges are resolved earlier, and the committee takes or explicitly declines the requested decision.
Change the production workflow, not only the scorecard
A rubric applied five minutes before the meeting becomes another compliance box. Put it into the work sequence.
1. Open with a decision brief. Before anyone prompts a model, record the client question, decision-maker, deadline, available options and minimum evidence. If the question cannot be written, the team is not ready to generate pages.
2. Build a source pack. Give the system approved documents, data extracts and definitions. Require citations to a source location, not a generic bibliography. A fluent paragraph without provenance remains an unverified claim.
3. Label content by function. Each section should be evidence, inference, option, recommendation, uncertainty or action. This makes duplicated narrative and unsupported jumps easier to spot.
4. Generate alternatives before layouts. Ask AI to challenge assumptions, identify missing evidence and compare options before asking it to design a storyline. Early polish creates attachment to weak analysis.
5. Review the decision chain first. A senior reviewer checks the eight DDS dimensions before fonts, page order and wording. Specialists inspect the evidence relevant to their domain; they do not need to reread every generated sentence.
6. Capture meeting outcomes. Record whether the client accepted, rejected, deferred or reframed the decision, plus the reason. “Deck delivered” is not an outcome.
This is where Lukas Hertig’s operator background matters. Across more than 20 years in infrastructure and hosting, scaling a business to €240 million ARR, a €1.5 billion exit and more than 15 acquisitions, the recurring lesson is that complex work needs an observable handoff. Boards do not act because a team manufactured more material. They act when evidence, consequence, authority and ownership meet in one traceable chain.
A 30-day proof path
Do not deploy DDS across the firm on day one. Test it on one recurring deliverable: a monthly performance pack, diligence readout, audit recommendation, tax-planning memo or steering-committee deck.
Days 1–5: Establish the baseline
Select three recently completed deliverables of the same type. For each, record:
- material decisions requested;
- DDS state by decision;
- client-facing pages per decision;
- senior review minutes;
- review rounds;
- evidence corrections after senior review; and
- meeting outcome: accepted, rejected, deferred or reframed.
Do not retrofit flattering scores. Missing records are themselves a finding.
Days 6–10: Install the decision brief
Choose the next live deliverable. Name the decision-maker, deadline and decision before drafting. Create the source pack and assign an owner for every material evidence block. Mark facts, inferences and recommendations separately.
Days 11–20: Run the new review gate
Draft with AI if useful, but prevent layout work until the chain is reviewable. Hold a 20-minute chain review. Any zero in verified evidence, decision requested or action owner returns the work to the author. Log review time and all material corrections.
Days 21–27: Use it in the client meeting
Put the requested decision on the agenda, not only in the final slide. Capture questions against the relevant DDS dimension. If the client disputes a baseline, that is an evidence failure. If the client agrees with the analysis but cannot act, inspect authority and ownership.
Days 28–30: Decide whether to scale
Compare the pilot with the baseline. Scale to one more team only if the deliverable required no more senior review time, had fewer late evidence corrections, and produced a clearer recorded decision or next action. Stop and redesign if page count falls but review time rises, evidence disputes move into the meeting, or teams game the rubric by splitting one vague recommendation into several “decisions.”
The point of the 30 days is not to prove a universal ROI percentage. It is to show whether a specific delivery workflow converts AI speed into cleaner professional judgment.
The new unit of professional output
AI will keep making first drafts, charts and layouts cheaper. Firms can respond by increasing volume, then spending partner time cleaning it up. Or they can change the unit they manage.
Do not ask how many slides the team produced today.
Ask:
- What client decision does this deliverable support?
- Can the client trace the evidence and reasoning?
- What remains uncertain?
- Who has the authority to decide?
- Who acts next?
- How much scarce review time did clarity require?
A compact deck with one complete decision chain can be more valuable than 60 polished pages. The DDS will not make the judgment for you. It will show where judgment is missing.
If you want to test the Decision-Density Score on one live client workflow, Book a 30-minute strategy call.
