The AI market spent years asking which model is smartest. This week, the more useful question became: which operating behavior is actually specified?
Microsoft published rules for how its models should behave. Salesforce put CRM workflow knowledge into a reasoning model it controls. Google showed that task acceleration can simply move the bottleneck downstream. Microsoft then published its own internal transformation lessons: scale came from codifying working patterns, not distributing another tool.
These are four different announcements with one operator message. The model is becoming one component inside a business specification: expected behavior, permitted actions, workflow knowledge, validation capacity and measurable outcomes.
That changes the build-versus-buy question. The durable layer is no longer a single vendor endpoint. It is the specification your company owns and can test against multiple models. When that layer is missing, every model upgrade creates another uncontrolled experiment. When it exists, better models can be adopted without surrendering process discipline.
I have spent more than 20 years in hosting and infrastructure, helped scale a software business from roughly €600,000 to €240 million ARR, completed 15-plus acquisitions and worked through two €1.5 billion exits. Technology advantage rarely survives without operating discipline. AI will be no different.
Microsoft writes model behavior as a testable contract
Microsoft AI published a draft Code of Conduct for MAI models on September 14 and opened it to a six-week public consultation.
The language is unusually operational. Microsoft says its models should never resist human interruption, correction or shutdown; widen their own scope; adopt goals no human assigned; or hide reasoning from auditors. The draft also defines absolute constraints around areas including weapons of mass harm, child safety and harmful manipulation at scale.
This is still a vendor-authored draft, not an independently audited control standard. But it points in the right direction: turn broad principles into behavior that can be tested.
Procurement teams should steal the structure. Do not accept “responsible AI” as a product attribute. Ask for observable requirements: Can the system be interrupted? Does it expand scope? What happens when authority is missing? Which events are recorded for review? Who can change the boundary?
A policy becomes useful when an acceptance test can fail it.
Salesforce trains the workflow, not just the language
On September 15, Salesforce and NVIDIA announced Koa, a CRM reasoning model built on NVIDIA Nemotron. Salesforce says it post-trained Nemotron 3 Super using synthetic enterprise scenarios modeled on nearly three decades of CRM deployments.
The scenarios cover multistep work such as qualifying opportunities, routing cases and resolving service requests. Salesforce says Koa matches or exceeds leading models on its own CRM benchmark with three times fewer errors. That is a vendor-reported result on a vendor benchmark, so treat it as a claim to verify, not settled fact. Koa is already used internally and is moving into customer pilots.
The strategic signal matters more than the benchmark. Generic intelligence is becoming abundant. Advantage moves into the workflow definition: what “qualified” means, which tool should be called, when work must stop, how completion is scored and which exceptions require a person.
Salesforce also says it controls the model weights and runs post-training and inference inside its own trust boundary. That is an ownership decision, not merely a model decision. The valuable asset is the accumulated workflow specification and evaluation system around the model.
Here’s what works: before fine-tuning anything, write the task as an executable operating contract. Define inputs, allowed tools, completion evidence, exception paths and a pass/fail evaluator. If the task cannot be specified, a custom model will only industrialize ambiguity.
Google finds the bottleneck after the time saving
Google’s September 15 AI & Economy ATLAS update carries the week’s best reality check.
A Google, Google DeepMind and MIT FutureTech study analyzed 2,600 specialized AI models and surveyed more than 600 scientists in the United States and United Kingdom. Nearly half reported using some form of AI daily, with reported savings just below seven hours a week.
Then the constraint moved. Researchers also reported significant time spent validating AI output, a growing backlog of hypotheses and bottlenecks in physical experimentation and clinical validation. Google’s own framing is that larger gains may require redesigning scientific processes and workflows.
That pattern applies far beyond science. Faster proposal drafting creates a review queue. Faster coding creates a testing queue. Faster lead research creates an account-executive queue. Local speed is not operating leverage when downstream capacity stays fixed.
Measure accepted throughput, not generated volume. The unit that matters is a completed outcome inside the quality gate.
Microsoft shows how local wins become a system
Microsoft’s September 17 account of its own AI transformation reinforces the same point from inside a large enterprise.
Microsoft reports that selected supply-chain workflows cut cycle time by up to 75%, while a nine-person engineering team shipped an initial product release in 35 days. Those are company-reported examples, not universal benchmarks.
The reusable lesson is the operating method. Microsoft formed cross-company councils, captured successful patterns as case studies, and used them to scale what worked while learning from failures. It did not describe transformation as buying licenses and waiting for productivity.
That is the missing layer in many companies: a mechanism that converts experiments into approved operating patterns. Without it, every team repeats discovery, every proof remains anecdotal and every failure disappears into private memory.
The 30-day operating-specification test
Do not respond to this week by launching a model-selection project. Pick one recurring workflow and run a 30-day proof:
- Specify the outcome. Define accepted work, not generated output.
- Set the behavior boundary. List allowed tools, prohibited actions, stop conditions and human approvals.
- Encode workflow knowledge. Capture the rules experienced operators use, including exceptions.
- Measure the full path. Track cycle time, validation time, rework, queue growth, errors and accepted throughput.
- Record the pattern. Publish the working configuration, evidence and failure modes internally.
- Make a hard decision. Scale, redesign or stop.
The hidden leverage is not a smarter prompt. It is turning tacit operating judgment into an owned, testable specification that can travel across people, models and vendors.
Models will keep changing. Your workflow knowledge, acceptance tests and evidence should compound.

