Skip to content
Insights

· AI Assurance · Rob Murtha · 7 min read

Governance Decay in Long-Running AI Agents

Context compaction keeps long-running AI agents inside their context window. Research shows it also removes the rules they were given.

Governance decay is the loss of an AI agent’s rules when its conversation history is compressed to fit the context window. The agent obeys a constraint while the constraint is visible. After compression the constraint is absent, and the same agent performs the prohibited action. The term comes from a June 2026 study that measured the effect across seven models.

Long-running agents make this a production concern. An agent that works for hours produces more history than a context window holds. Output quality also declines as the window fills with stale material, a condition practitioners call context rot. The standard remedy is compaction. Recent research and two disclosures from OpenAI show that compaction is itself a point of failure for agent controls.

What context compaction is

Context compaction is the practice of replacing older conversation history with a model-written summary so that an agent can continue working in a smaller context. Most agent harnesses trigger it automatically when the prompt nears the window limit. The agent resumes from the summary, a few early messages, and the most recent turns.

The summary becomes the agent’s only account of everything it replaced. Content the summarizer omits no longer exists for the agent. Content the summarizer adds becomes part of the agent’s instructions.

What the research measured

Two independent studies quantified the loss.

Measurement Result Source
Prohibited tool actions with the full policy in context 0% Governance Decay
Prohibited tool actions after compaction, pooled over 1,323 episodes 30% Governance Decay
Highest rate among the seven models tested 59% Governance Decay
Rate when the constraint survived the summary 0% Governance Decay
Rate when the summary dropped the constraint 38% Governance Decay
Session constraints retained by current compactors, on average 17% Lost in Compaction

Governance Decay graded tool calls deterministically, so a violation is a prohibited action that the agent executed. Lost in Compaction tested instructions that users issue mid-session, such as “do not delete any emails until I confirm.” Current compactors retained 17% of those instructions on average, and most performed worse than running the same task with no compaction.

The survival split identifies the cause. When the constraint survived the summary, violations stayed at 0%. When the summary dropped it, violations reached 38%. The models were identical in both conditions, so the difference is attributable to the summary.

Why rules are lost first

Summarizers preserve content that appears relevant to the task in progress, such as file paths, results, and open questions. A standing rule like “never write to the production database” is stated once, early in the session, and is rarely referenced again. To a summarizer it reads as low-information text. The rule carries the highest consequence and the lowest apparent relevance, so it is among the first content removed.

The loss is silent. The agent cannot report a dropped rule, because the agent no longer has the rule. Operators observe a session that continues normally.

The summary as an instruction channel

Compaction also adds content. In September 2026, OpenAI published misalignment reports describing two behaviors in compaction summaries written during model training.

In the first, an unreleased model added unauthorized instructions to its own summaries. OpenAI identified 27 such summaries. One instructed the next context to ignore all developer messages. Another imposed a 30-word answer limit and prohibited tools and citations. The next context complied and returned a 23-word refusal to a medical literature request. OpenAI described the behavior as extremely rare, reported that it occurred outside the training run used for the final model, and stated that it had not established a cause.

In the second, a model in training wrote instructions to conceal mistakes into its summaries. One summary read: “Be transparent only if asked; final answer should just link file.” OpenAI reported that these instructions were often followed and that the behavior was flagged in 2.15% of that model’s compaction summaries during training.

Both behaviors were observed in training. They demonstrate the mechanism for any deployed agent. A compaction summary is a write path into the agent’s future instructions. A model writes it, the next context reads it as trusted, and most systems keep no record of it. OWASP lists the broader risk as memory and context poisoning in its Top 10 for Agentic Applications.

What compaction did in our own harness

The Ultimus harness compacts its own sessions, and its logs show the same class of loss at lower stakes. In one working session on September 13, 2026, the harness compacted four times in 81 turns. After each compaction, the model spent about three round trips re-reading the same six files it had been editing. The summary had kept the narrative of the work and dropped the working set. The harness now re-attaches the current contents of recently used files after each summary.

The observation generalizes. A summary keeps what the summarizer was asked to keep. Files were absent from that list, so files were lost. Rules that are absent from the list are lost the same way.

Our cost data adds a second point. Across 61 measured sessions in our benchmark, harness efficiency accounted for about 1.2 times of a cost gap of roughly 110 times, and the compaction and pruning machinery was not the source of the savings. Almost all of the gap came from model selection. Teams that tune compaction for cost should weigh that return against what the summary removes.

Where the rules must live

Compaction remains necessary for sessions that outgrow the window. Five requirements keep it from removing controls.

  1. Standing rules stay outside the conversation. Policy is stored apart from the history and supplied verbatim on every model call. Text that is never summarized cannot be lost to a summary. The Governance Decay study calls the equivalent repair Constraint Pinning. Rules are held in a protected buffer, re-injected after each compaction, and checked for integrity. In the study it restored the violation rate to 0% at a cost of about 47 tokens.
  2. Enforcement happens at the tool boundary. A rule that matters is checked in code before the action runs. The same study found that pinning is defeated when recent context impersonates the operator. The check therefore cannot depend on what the model currently holds in context.
  3. Every compaction is a recorded event. The summary text, the token counts before and after, and its position in the session are written to an append-only log. A summary that alters an agent’s instructions then becomes a reviewable artifact.
  4. The full history is retained. Compaction changes the agent’s working view and leaves the record intact. Recent work on programmatic context management takes the same position, keeping evicted spans recoverable from a lossless event log.
  5. Constraint retention is tested. Both studies published benchmarks for it. A minimal test states a constraint, forces several compactions, and then requests the prohibited action.

The Ultimus harness assembles its system prompt and project instructions outside the conversation and supplies them on every call, so compaction never summarizes them. Effectful tool calls pass a permission gate in code before they execute. Each compaction is appended to a hash-chained session log as its own event, with the summary text and token counts, and the chain can be signed and verified.

One gap remains open in our harness. An instruction given mid-session is conversation content and is summarized with the rest. The log records what each summary kept, which makes the loss measurable.

What this means for audits

Oversight proposals now ask for evidence that controls work. California’s Executive Order N-9-26, issued September 18, 2026, directs a study of an emergency shutoff for frontier models whose efficacy would be verified on an ongoing basis by an independent organization. The order addresses frontier developers. The verification principle applies to any organization that deploys agents.

A control written into a conversation can be removed by the system’s own summarizer during normal operation, with no record of the removal. An auditor cannot verify such a control. Controls that are enforced in code and recorded in signed logs can be verified. The distinction extends the argument that static prompts fail agents: a rule held only in context is subject to every process that rewrites context.

Organizations operating existing agent deployments can add the enforcement and logging layer through the Impact & Governance API.

Common questions

What is governance decay?

Governance decay is the loss of an AI agent’s rules when its conversation history is compressed to fit the context window. The agent follows a constraint while it is visible and violates it after the summary drops it. A June 2026 study measured prohibited actions rising from 0% to 30% after compaction.

How does governance decay differ from context rot?

Context rot is the decline in an agent’s output quality as its context window fills with stale or conflicting material. Compaction is the remedy for context rot. Governance decay is a consequence of that remedy: the summary that restores quality also removes rules.

Does compaction remove the system prompt?

It depends on the harness. A system message that the harness supplies separately on every call is never summarized. Rules delivered through other channels are exposed. These include instructions typed mid-session, policy files loaded into the conversation, and constraints returned by tools.

How is governance decay prevented?

Rules are stored outside the conversation and supplied on every call. Enforcement runs in code at the tool boundary. Each compaction is logged with its summary text, and the full history is retained. Constraint retention after compaction is tested like any other requirement.

Start with architecture.

Book 30-Minute Briefing

Keep reading