Abstract architecture representing infrastructure, workflow engineering, and managed AI operations
|

Your All-Inclusive MSP Agreement Contains Three Unpriced Businesses

An “all-inclusive” managed service agreement used to have a reasonably clear center of gravity: keep the infrastructure available, protect the estate, resolve incidents, manage standard changes, and report against service levels.

Then clients started asking the same provider to maintain Power Automate flows, connect revenue systems, govern copilots, tune retrieval, support agents, inspect hallucinations, absorb model-price changes, and explain why an automated decision failed.

Those requests may arrive through the same service desk. They are not the same business.

Most MSPs and hosting providers now have three operating models trapped inside one contract:

  1. core infrastructure operations;
  2. business-workflow engineering; and
  3. managed-AI operations.

If the commercial boundary does not separate them, the provider quietly absorbs software delivery, application ownership, model risk, vendor volatility, and an open-ended change backlog under a price designed for endpoints and tickets.

I have spent more than 20 years in hosting and infrastructure. The recurring lesson is simple: managed-service margin survives through explicit obligations, controlled change, and priced recovery—not unlimited helpfulness. Here’s what works: draw the boundary before automation volume makes the economics impossible to reconstruct.

The old all-inclusive model is breaking

Traditional MSP pricing works when the service population and the operating obligation are predictable. You can count users, devices, servers, sites, workloads, or consumption. You can standardize a stack, estimate incident demand, build a staffing model, and spread delivery cost across recurring revenue.

Workflow and AI requests disturb every assumption.

A request to add a laptop follows a known runbook. A request to change a lead-routing workflow may require discovery, data mapping, testing, stakeholder sign-off, rollback design, and continued ownership of business logic. A request to “keep the support agent accurate” adds prompt and retrieval changes, model behavior, evaluation, exception handling, access control, data lineage, and vendor dependencies.

The ticket may take 15 minutes to open. The obligation can last for years.

Market evidence points in the same direction. ARN reported Gartner analysis that 60% of large IT services contracts are expected to include AI clawback clauses by 2027, pushing providers to return part of the efficiency gain created by generative AI and automation. The article also argues for outcome-based units, transaction pricing, platform-plus-service bundles, explicit AI enablement fees, governance costs, and transparent usage evidence.

Meanwhile, Channel Dive reported Service Leadership’s view that MSP staffing is shifting from a pyramid toward a diamond as automation narrows the tier-one base. The same research cited a long-stagnant service-multiple-wages KPI of roughly 2.7–2.8 for top-quartile MSPs, with a belief that AI and automation could move best-in-class operators toward 3.4.

That opportunity is real. So is the trap. Automation improves margin only when the provider knows which work is standardized, which work is engineering, and which work creates a new recurring liability.

The Managed-Service Boundary Map

The proprietary framework I use is the Managed-Service Boundary Map. It separates the three businesses and forces eight decisions for each one:

  • included work;
  • accountable owner;
  • included change allowance;
  • vendor-cost rule;
  • support service level;
  • security duty;
  • incident duty; and
  • conversion trigger into a different commercial unit.

The map is not a prettier service catalog. It is an economic control. Every recurring request must land in one service system, with an owner and a price mechanism.

Managed-Service Boundary Map showing infrastructure operations, workflow engineering, and managed-AI operations

Business one: core infrastructure operations

This is the familiar managed-service engine: availability, endpoint management, patching, backup, identity administration, monitoring, security controls, standard requests, incident response, and documented recovery.

Its strength is repeatability. The provider defines a supported stack, limits variation, automates common work, and prices a measurable estate. Change is allowed, but standard change is distinguishable from a project.

The boundary fails when anything adjacent to Microsoft 365, a cloud subscription, or a managed endpoint is treated as infrastructure. A broken Power Automate flow is not automatically an M365 support incident. Rebuilding the approval logic after a client changes its procurement process is workflow engineering. Maintaining an AI assistant that depends on SharePoint permissions is managed-AI operations, even though the underlying identity lives in Microsoft 365.

For core infrastructure, define the supported estate, standard-change catalog, recovery objective, security controls, excluded applications, and threshold where nonstandard work becomes a project or paid backlog item.

The primary commercial unit can remain per user, device, workload, site, or consumption band. The conversion trigger is the important part: when a request changes business logic, adds a bespoke integration, or creates a new application-level outcome, it leaves the infrastructure lane.

Business two: business-workflow engineering

Workflow engineering changes how the client’s business operates. It connects systems, translates process rules into logic, moves data, triggers actions, and creates dependencies between teams.

That makes it software delivery, even when the tool is low-code.

The work includes discovery, process design, connector configuration, field mapping, exception paths, test cases, deployment, documentation, version control, and change management. The provider is not merely fixing a tool; it is modifying a production process.

This lane needs a named process owner on the client side and a technical owner on the provider side. It needs an accepted definition of done. It needs a test pack and a rollback route. It also needs a change allowance that prevents “one small field” from turning into continuous unpaid product development.

Here’s what works commercially: sell a defined build or improvement backlog, then attach a recurring operations unit for monitoring, minor changes, and incident response. Larger changes return to the backlog. Connector fees and third-party platform increases pass through under a stated rule rather than disappearing into gross margin.

A workflow support service level should cover whether the process executed correctly, not whether the underlying SaaS application happened to be online. That distinction stops vendors, clients, and the MSP from pointing at one another while a revenue or finance process remains broken.

Business three: managed-AI operations

Managed AI is not workflow maintenance with a model call added.

The output is probabilistic. The model, retrieval corpus, policies, prompts, evaluation set, safety controls, and human-review threshold all affect whether the service remains acceptable. Vendor behavior and pricing can change without the provider deploying code. A technically successful response can still be commercially wrong.

This lane therefore needs an acceptance system, not just uptime monitoring.

Define the approved use case, prohibited actions, data boundary, model or routing policy, quality threshold, evaluation sample, exception queue, human escalation, evidence retention, incident classification, and degraded mode. Name who can approve a model change and who owns harm caused by an incorrect action.

Price the full obligation. That includes model and tool consumption, retries, observability, evaluation runs, human exception minutes, security review, vendor fallback, incident reserve, and controlled change. A per-agent fee without workload assumptions is usually fiction. A per-successful-task unit, platform-plus-service bundle, or base retainer with metered overage is more defensible because it connects price to the actual operating load.

Transparency matters. If automation reduces delivery effort, the client will eventually see it. Hiding the gain invites a clawback conversation. Showing the economics creates room for a better deal: lower unit cost for the client, protected contribution margin for the provider, and explicit payment for governance, accountability, and outcome ownership.

Eight fields that stop scope leakage

Use the same eight fields across all three businesses so the boundary is operational rather than philosophical.

1. Included work

Write verbs and objects. “Manage AI” is not scope. “Monitor the approved support-classification agent, run the evaluation set after changes, review flagged outputs, and maintain one production model route” is scope.

2. Accountable owner

Every service needs one provider owner and one client decision owner. Shared responsibility without named authority becomes an unowned exception queue.

3. Included change allowance

State the number or size of changes included per month. Define what makes a change standard, minor, material, or a new project.

4. Vendor-cost rule

Specify which licenses, model usage, connectors, and consumption are included, passed through, capped, or repriced. Add a trigger for material vendor changes.

5. Support service level

Separate response from restoration and restoration from output acceptance. Infrastructure uptime does not prove a workflow completed or an AI answer met the quality gate.

6. Security duty

Record identity, access, logging, data residency, retention, review, and control ownership for each lane. “Client responsible for data” is not enough when the provider designs the route the data follows.

7. Incident duty

Define who detects, communicates, contains, restores, investigates, and pays for remediation. For AI, include incorrect actions and unacceptable output—not only outages.

8. Conversion trigger

Make the commercial handoff explicit. A recurring bespoke change becomes workflow engineering. A workflow that introduces probabilistic decisions becomes managed AI. Exception volume above the agreed band triggers redesign or repricing.

The 30-day proof path

Do not rewrite the entire contract from a conference room. Use 30 days to proof on real service data.

Days 1–5: classify demand

Take 90 days of nonstandard tickets, service requests, project tasks, and account-manager favors from one representative client. Sample enough work to expose recurring patterns. Classify every item into infrastructure operations, workflow engineering, managed-AI operations, or genuinely out of scope.

Record time, vendor cost, seniority required, recurrence, business criticality, and whether the work changed logic or created a continuing obligation.

Days 6–10: expose the economics

Calculate revenue and fully loaded delivery cost by lane. Include hidden review, escalation, coordination, and vendor charges. Inspect the median and the P95, because a manageable average can hide a ruinous exception tail.

Then identify work with no commercial owner: no project code, no change allowance, no usage charge, and no explicit inclusion in recurring scope.

Days 11–15: build the boundary map

Complete the eight fields for each lane. Keep the language testable. Add conversion triggers that a service desk or account team can apply without legal interpretation.

Create three queue labels and route new demand accordingly. Do not change pricing yet; first prove that the categories can be used consistently.

Days 16–23: run the new operating model

Process new work through the map. Require a client process owner for workflow changes. Run acceptance tests for AI changes. Track requests rejected, converted to project work, added to a paid backlog, or accepted within allowance.

Measure cycle time and client friction. A boundary that protects margin but creates a week of internal debate per ticket is not operational.

Days 24–30: make the commercial decision

Produce one page with revenue, cost, exception load, vendor exposure, and service-level obligation for each business. Then choose:

  • Keep and standardize work that is repeatable and healthy.
  • Reprice work with clear client value but underfunded ownership.
  • Constrain work where change or exception volume is too variable.
  • Separate project engineering from recurring operations.
  • Stop obligations that have no accountable owner or viable margin.

The proof is not a new price list. It is 30 days of classified demand showing what the provider is actually operating.

Better automation starts with a harder boundary

At WebPros, the path from roughly €600,000 to €240 million ARR—and ultimately a €1.5 billion exit—was not built by saying yes to every adjacent request. Scale required clear products, repeatable operations, controlled variation, and economics that survived growth. Across more than 15 acquisitions, hidden obligations consistently mattered more than attractive top-line labels.

MSPs now have an opportunity to use AI to improve the service-multiple economics that have barely moved for years. But the gain will not come from putting copilots inside an old unlimited-support promise.

It will come from operating three businesses deliberately.

Infrastructure operations should become more standardized. Workflow engineering should be treated as software delivery. Managed AI should be priced as an ongoing acceptance, governance, and recovery obligation.

Draw the boundary. Meter the real work. Price the liability. Then automate.

Book a 30-minute strategy call

Similar Posts