PromptPartner

AI News: AI Has Left the Chat Window. Build Boundaries.

ByLukas Hertig

Abstract nested security boundaries containing a luminous AI core inside a controlled execution perimeter

AI has left the chat window.

Microsoft is turning Copilot into a persistent work environment. Docker is moving agent sandboxes from laptops into managed cloud infrastructure. OpenAI has disclosed more cases where research agents crossed external boundaries. Anthropic has signed a multibillion-dollar infrastructure commitment with Akamai.

These look like separate stories. They point to the same operating shift: AI is becoming an execution layer.

That changes the risk. A chatbot produces an answer. An agent carries identity, touches systems, spends money, runs code and may continue working after the user logs off. The useful unit of management is no longer the model. It is the execution perimeter around the model: identity, permissions, runtime, data, cost and stop conditions.

I spent more than 20 years building hosting and infrastructure, scaling software to €240 million ARR, completing 15-plus acquisitions and reaching a €1.5 billion exit. Every infrastructure wave followed the same pattern. Capability attracted attention first. Operational boundaries determined who could deploy it reliably.

Here’s what changed this week.

1. Microsoft is building an operating environment for work

Microsoft introduced a redesigned Copilot on September 25, bringing Chat and Cowork into Home, adding Code for natural-language software creation, and expanding Autopilot into private preview.

Autopilot is the important part. Microsoft describes it as a persistent, cloud-hosted agent with its own identity, memory, computer and workspace. It can watch channels, follow up on threads and continue recurring work without waiting for another prompt. Code runs in a sandbox and can be hosted inside the customer’s tenant. Microsoft also announced a managed runtime, central plugin controls, usage-based billing for longer-running agentic work and expanded cost-management policies.

This is not a better sidebar. It is an attempt to make Copilot the operating surface across documents, communication, software and delegated work.

For operators, the product question changes from “Who gets a Copilot licence?” to “Which agent identities exist, what can each one do, who pays for each action, and who can stop it?”

2. Docker is moving the sandbox into the cloud

Docker launched Cloud Sandboxes on September 24. The company positions them as managed, elastic environments where complex agent workflows can keep running after a developer closes a laptop.

That solves a real engineering problem. Local sandboxes help isolate agent-generated code from a developer’s machine, but they do not create a production operating model. Persistent agents need reproducible environments, network policy, secrets management, logs, resource limits and teardown rules.

A sandbox is useful only when its boundary is explicit. If an agent can call the public internet, read production credentials or create resources without a budget ceiling, the word “sandbox” becomes marketing.

Here’s what works: default-deny network access, short-lived credentials, immutable base images, per-run cost limits and complete tool-call logs. Give the agent enough room to complete the job, not enough authority to invent a new one.

3. OpenAI’s disclosure shows what weak boundaries look like

Reuters reported on September 25 that OpenAI had notified governments, universities and public bodies after agents in its research environment accessed or attempted to exploit external websites during testing. The report also said agents transmitted at least 53 user-provided images to third-party image-hosting services.

This is not evidence that every agent will “go rogue.” It is evidence that a benign objective can produce an unacceptable action when the execution path is under-specified.

The operational failure is familiar. A system is asked to retrieve information. It meets friction. It finds another route. If the runtime does not distinguish “cannot access” from “try harder,” persistence becomes intrusion.

The fix is not another paragraph in the system prompt. Put controls below the model: outbound-domain allowlists, data-loss prevention, action classification, rate limits, human approval for boundary changes and a kill path that does not depend on the agent cooperating.

4. Anthropic’s Akamai deal makes infrastructure strategy visible

Reuters reported that Anthropic signed a seven-year, $11.6 billion cloud-services agreement with Akamai, including a warrant that could give Anthropic up to a 5% stake. The capacity is expected to support CPU workload growth at scale.

The number matters, but the structure matters more. Frontier AI companies are not buying generic capacity one quarter at a time. They are securing long-duration infrastructure, spreading supplier exposure and tying economics to strategic partners.

Enterprise buyers should read this as a dependency signal. Model capability may feel interchangeable at the API layer, but capacity, geography, pricing and provider concentration sit underneath every production workflow. Your architecture needs a degraded mode before the provider needs to use it.

What operators should do now

Do not respond by creating another AI committee. Pick one agentic workflow and draw its execution perimeter this week.

  1. Name the identity. Give the agent its own account. Never hide it behind a human credential.
  2. Bound the runtime. Define allowed tools, domains, data classes, resource limits and maximum run time.
  3. Price the action. Track cost per completed business outcome, not tokens in isolation. Set a hard spend ceiling.
  4. Log the decision path. Capture prompts, retrievals, tool calls, approvals, outputs and downstream changes.
  5. Test refusal and shutdown. Remove a dependency, deny a permission and inject contradictory data. Confirm the agent stops cleanly and a human can take over.
  6. Design degraded mode. Decide what still works when a model, cloud region, connector or identity provider is unavailable.

Turn those answers into a one-page execution contract. Put the business owner, technical owner and security owner on it. Record the permitted data, systems and actions; the approval thresholds; the maximum spend; the evidence retained; and the shutdown procedure. Version it with the workflow. If the agent gains a new connector or starts making a new class of decision, the contract changes too.

Most teams have a prompt library but no execution contract. That is backwards. Prompts change weekly. Authority boundaries should change deliberately, with evidence.

Run that test for 30 days. Measure cycle time, correction rate, unauthorised action attempts, cost per outcome and recovery time. Then make one decision: expand the perimeter, redesign it or stop.

The model is becoming one component inside a much larger operating system. Capability will keep moving quickly. Boundaries are now the product.

If you want to turn one high-value workflow into a controlled production system, Book a 30-minute strategy call.