Where the instruction disappears
Researchers at Penn State have put a number on something builders have felt for a while. Their evaluation suite, COMPINT, injects session rules — the ordinary side conditions a user types once, like 'confirm with me before making any changes' — and then measures how many are still in force after the assistant compacts its conversation history to free up room. On average, 17 per cent survive. Most of the tested compaction setups score worse than running the same task with no compaction at all, and that is the part that should stop a builder mid-scroll.[1]
The size of the drop is worth holding onto. With the whole uncompressed context present, compliance with the rule runs between 59 per cent and 71 per cent, an imperfect but genuine signal. After compaction, most compactors land close to the level measured when no rule was given at all. A compaction prompt written specifically to preserve user constraints still leaves retention under 40 per cent, while GPT-5.4-mini beats the uncompressed baseline in some scenarios, so the loss varies with the system doing the summarising.[1]
The mechanism behind the number looks architectural. A compactor is built to keep a task moving: the goal, the current state, the next steps. A session rule belongs to none of those; it constrains behaviour, and the summariser has no structural reason to carry it. That reading is an inference from the pattern rather than a measured cause, and a plausible alternative exists: session rules tend to be short, arrive early and never recur, so a plain length-and-recency heuristic could produce the same losses without any bias about what a rule is. Either way the practical consequence is identical. An agent told to ask before sending mail will, after compaction, send it.[1]
The fix sits beside the model
What the researchers propose is small enough to be interesting. A separate model built on Qwen3.5-9B runs alongside the compactor, reads every user input, collects session rules into its own list and appends that list to the summary. Retention goes above 90 per cent in all three tested scenarios: 95.6 per cent on agent trajectories, 95.1 per cent on long-horizon research and 90.3 per cent in multi-turn chat. It needs no training and no change to the compaction system, and both COMPINT and the extractor are published on GitHub.[1]
That shape matters more to me than the accuracy figure. The thing that closes the gap is a component a team can drop in, inspect, replace or remove — a part of the toolchain rather than a property of the weights. It continues a line I took in this column on 14 August, reading Anthropic's multi-agent study: the swarm's advantage and its collapse both came from whether the shared environment had been built, which made naming, arbitration and reputation the builder's load. Constraint retention is the same kind of load, one layer down, and it lands on the team assembling the stack.[1], [3]
AMD's numbers point at the same layer
AMD published a figure this week that fits the same shape from the industrial end. Agents fixing reported issues in Radeon Software eXperience resolved 6 per cent of them when the effort began in October 2025 and more than 75 per cent by June 2026. The senior vice president who wrote it up credits refining the objectives given to the agents and letting them explore several approaches against defined success criteria, instead of retraining the underlying models. This is the company's own account in an opinion piece, with no independent audit and no numbers on review time or defect rates, so it stays attributed vendor evidence.[2]
Put the two side by side and the common constraint is visible: in each case the binding limit was what the model had been told and how much of it survived, and in each case the repair happened outside the weights. That suggests a test with a date on it. If, by 31 October 2026, an agent framework ships a documented constraint-retention slot — user rules held where the compactor cannot touch them — and publishes a retention figure on a public suite such as COMPINT, then constraint handling will have moved from prompt craft into the toolchain. If no framework documents such a slot by then, the 17 per cent stands as a property teams have to work around one prompt at a time.[1], [2]