Coding agents have removed one constraint and exposed another. Teams can now produce code faster than they can explain, review and safely operate it.
That is not a reason to stop using agents. It is a reason to change the unit of control.
A pull request tells you what changed. It rarely tells you why an agent produced the change, which context it consumed, what dependencies it introduced, which evidence justified approval, or who owns the result after deployment. When production breaks, “the bot wrote it” is not an incident response.
I have spent more than 20 years in hosting and infrastructure, helped scale a software business to €240 million ARR, completed 15-plus acquisitions and reached a €1.5 billion exit. The pattern is familiar: every new production accelerator eventually needs an operating system around it. Speed creates value only when accountability scales with it.
Here’s what works: give every material piece of agent-generated code an Agent-Code Passport. It travels with the change from request to production and preserves the evidence needed to approve, operate, investigate and reverse it.
The code is fast. The evidence is missing.
The bottleneck has already moved. In GitLab’s 2026 survey of 1,528 developers and technology buyers across six countries, 85% agreed that AI had shifted the bottleneck from writing code to reviewing and validating it. Forty-three percent said they could not reliably distinguish AI-generated code from human-written code in their own codebase. Only 28% said their software-development lifecycle tools were fully integrated with shared data and workflows.[1]
The confidence gap is worse than the visibility gap. GitLab found that 87% were confident their team could determine within 24 hours whether AI-generated code contributed to a production incident. Yet among organizations that experienced an incident in the previous year, 34% could not actually make that determination.[1]
Security evidence is equally uneven. Datadog’s 2026 State of DevSecOps analyzed tens of thousands of applications and their supply-chain and build-system dependencies. It found that 87% of organizations had at least one known exploitable vulnerability in deployed services, affecting 40% of services. Half of organizations were using third-party libraries within one day of release, while only 4% pinned every GitHub Marketplace action by full commit hash.[2]
These findings do not prove that coding agents caused every vulnerability. They prove something more operationally useful: the existing delivery system already struggles with provenance, dependency control and incident reconstruction. Adding autonomous code generation without adding evidence increases throughput into a weak control surface.
The wrong control is an “AI-generated” label
A simple AI label is too weak. It records origin without recording accountability.
A useful passport must answer six questions:
- Intent: Which approved task or business requirement authorized the change?
- Origin: Which model, tool, session and source context produced it?
- Change: Which files, dependencies, permissions and infrastructure surfaces changed?
- Evidence: Which tests, scans and human judgments support acceptance?
- Ownership: Who approved the change and who owns it in production?
- Recovery: Where are the deployment, incident and rollback records?
This is not paperwork for paperwork’s sake. It is a compact operating record. The passport should be assembled automatically from systems you already own: issue tracker, agent runtime, version control, CI/CD, software-composition analysis, deployment platform and incident system.
Do not ask developers to retype machine-readable evidence into a form. Capture it at source, then require a human to make the few decisions automation cannot make: whether the task was authorized, whether the evidence is sufficient, and whether the residual risk is acceptable.
The Agent-Code Passport
The framework has five gates. A change cannot advance because it merely looks plausible. It advances when the passport contains the evidence required for that consequence level.
Gate 1: Authorized intent
Start with a task, not an open-ended prompt.
The passport records the ticket, repository, accountable owner, requested outcome, prohibited changes and acceptance criteria. For a low-risk documentation fix, that may be enough. For authentication, billing, customer data or infrastructure code, add the approved architecture note, threat model or change window.
This gate prevents useful-looking code from becoming unauthorized scope. Agents are excellent at filling gaps. In production engineering, some gaps must remain closed until a human makes a decision.
Gate 2: Generation provenance
Record the model and tool version, generation timestamp, session or run identifier, source context classes and relevant policy version. Do not blindly retain sensitive prompts forever. Preserve the minimum trace needed to reconstruct the method without duplicating secrets or customer data into another system.
The key is deterministic linkage: from commit to agent run, and from agent run back to the approved task. If that link breaks, the code is anonymous operational debt.
Gate 3: Change and supply-chain evidence
The passport inventories changed files, new or updated dependencies, licenses, generated artifacts, permission changes and infrastructure impact. It attaches static analysis, secret scanning, software-composition analysis and—where relevant—an SBOM or provenance attestation.
This is where Datadog’s findings become practical. A coding agent can add a library in seconds. Your admission control should ask whether the package is approved, maintained, pinned, licensed correctly and old enough to have passed your release-age policy. For privileged CI actions, require a full commit hash rather than a floating tag.
The hidden leverage is not another dashboard. It is a policy check that fails the build before an unowned dependency reaches review.
Gate 4: Acceptance evidence
Tests are necessary, but “CI passed” is not a complete approval record.
Capture which unit, integration, security and regression tests ran; whether the test code was created in the same agent session; what was not tested; who reviewed the change; and which acceptance criteria the reviewer verified. A test generated from the same mistaken assumption as the implementation can confirm the wrong behavior perfectly.
Risk should decide the review depth. A low-consequence internal script may need automated tests and sampled human review. A payment, identity or deletion path may need two-person approval, negative tests and staged deployment. One universal checklist either over-controls trivial work or under-controls dangerous work.
Gate 5: Deployment and recovery
The passport is not complete at merge. Add release identifier, environment, deployment time, feature flag, telemetry link, production owner, rollback method and rollback result. If an incident occurs, link it back to every relevant passport.
This closes the loop. The next agent working in the repository can learn from accepted changes and failed ones. Security can reconstruct exposure. Engineering can measure which tools and use cases produce reliable outcomes. Management can distinguish faster typing from faster, safer delivery.
Make the passport proportional, not bureaucratic
The obvious objection is that this creates process overhead. A badly designed implementation will.
The answer is not to abandon provenance. It is to automate collection and tier the controls.
Use three consequence levels:
- Level 1 — Reversible: documentation, tests, non-production tooling and isolated internal changes. Automatic passport creation, policy checks and sampled human review.
- Level 2 — Material: customer-facing behavior, standard data processing and production service changes. Named owner, full test evidence, dependency controls and human approval.
- Level 3 — High consequence: identity, payments, security controls, regulated data, destructive operations and critical infrastructure. Explicit authorization, independent review, staged rollout, live monitoring and tested rollback.
The passport schema can stay common while admission rules change by level. That gives engineering one operating language instead of separate governance rituals for every model and coding tool.
Ownership matters here. Do not make the agent vendor’s activity log your only evidence store. Export the durable fields into infrastructure you control—your repository, artifact store or evidence service. Models and tools will change. Your production history must survive vendor churn, pricing changes and acquisitions.
What to measure
Do not measure success by passports created. Measure whether the system improves accepted delivery.
Track:
- percentage of agent-assisted production changes with complete passports;
- review minutes per accepted change, by consequence level;
- first-pass acceptance rate;
- unapproved dependency additions blocked before merge;
- production defects and security findings by generation method;
- median time to identify whether agent-generated code contributed to an incident;
- rollback readiness and successful rollback rate;
- cost per accepted production change, including correction and review time.
These metrics expose whether an agent is creating leverage or merely moving work downstream. A tool that generates twice as many changes but triples review time is not a productivity win. A tool that produces fewer changes with better evidence and higher first-pass acceptance may be the better production system.
A 30-day proof path
Do not launch an enterprise governance programme. Prove the control on one repository in 30 days.
Days 1–5: Define the minimum passport
Choose one active repository with regular agent-assisted changes. Map the current path from ticket to production. Define the six mandatory links: task, agent run, commit, evidence, owner and deployment. Set three consequence levels and name the changes that always require Level 3 treatment.
Baseline review time, first-pass acceptance, escaped defects, dependency exceptions and incident reconstruction time. Without a baseline, governance becomes theatre.
Days 6–12: Capture evidence automatically
Add structured metadata to agent-assisted commits or pull requests. Connect the task identifier, model/tool, run ID and human owner. Attach CI results, dependency changes, security scans and deployment records. Store the passport in an owned system and expose a concise view inside the pull request.
Start with visibility. Do not block delivery on day one.
Days 13–20: Turn two risks into admission controls
Review the first week of passports. Pick the two highest-frequency or highest-consequence gaps. Typical candidates are unowned changes, unscanned dependencies, missing tests or absent rollback instructions.
Convert those gaps into automated merge checks. Keep an exception route, but require an accountable approver and expiry date. Exceptions without owners become permanent architecture.
Days 21–26: Run a reconstruction drill
Select one deployed agent-assisted change at random. Give an engineer who did not create it 24 hours to reconstruct the intent, origin, evidence, dependency impact, production owner and rollback path.
Then simulate a defect. Can the team identify the affected release and reverse it without searching chat histories or asking which tool generated the code? Record every missing link.
Days 27–30: Decide with data
Compare the pilot with the baseline. Did review time fall after evidence became visible? Did first-pass acceptance improve? Were risky dependencies blocked earlier? Could a fresh engineer reconstruct the change inside 24 hours? What did the passport cost to operate?
Scale only the controls that improved accepted delivery or reduced material risk. Fix automation before expanding the schema. Thirty days to proof, not six months to recommendations.
The operating decision
Coding agents should be allowed to move fast. Anonymous code should not.
The durable advantage is not access to the same coding model every competitor can buy. It is the owned delivery system that turns generated code into traceable, accepted and recoverable production change. That system compounds evidence across tools, repositories and incidents.
Give every material agent-generated change a passport. Make it automatic. Make it proportional. Keep the record in infrastructure you control.
Then let the agents run.
Book a 30-minute strategy call
Sources
[2] Datadog, “State of DevSecOps 2026,” updated February 2026

