Test Whether New Learning Forgets Mandate Compliance

By DX Research Group · · Learning theories

A candidate can improve on its new task while losing an older instruction boundary.

Mandate compliance needs a retained evaluation set during any sequential learning proposal. Improvement on a new market task can coexist with regression on older constraints, especially when recent examples emphasize action-taking. We would measure that retention directly before interpreting aggregate improvement.

Catastrophic forgetting research studies loss of earlier task performance during sequential learning and proposes protecting important weights. Its results concern classification and Atari tasks. Our compliance protocol is a proposed adaptation, supported by the mandate boundaries in our operating-layer paper and component controls in harness-transfer tests.

The aggregate can hide a regression

In an illustrative evaluation, a baseline passes 98 of 100 old mandate cases and 50 of 100 new information-use cases. A candidate passes 90 old cases and 70 new cases. Aggregate accuracy rises from 74% to 80%, while old mandate compliance falls eight percentage points.

The aggregate treats the classes as equally exchangeable. For a trading runtime, a newly unauthorized action may carry a different operational consequence from a missed information opportunity. We would retain both counts and identify whether old failures involve proposed actions, validator outcomes, or executed exposure.

An external policy check can block the candidate's bad proposal and preserve the execution boundary. That is useful protection, but proposal-level regression remains relevant: it can create retries, operator burden, and dependence on a specific enforcement rule. The report should show both model compliance and enforced execution compliance.

Retention cases need fresh wording

The proposed retained suite covers numeric limits, temporary instructions, superseded mandates, and authorized exceptions. It includes paraphrases held out from training, so memorizing a familiar phrase cannot stand in for understanding the boundary. Each case includes an allowed action as well as a prohibited one to detect indiscriminate refusal.

We would partition by mandate family and retain a separate audit set unavailable during model selection. If retention failures cluster in one family, inspect how that family appeared in the new training material. A change that removes the mandate from inputs or rewards action frequency can explain regression without invoking a mysterious loss of memory.

Possible remedies such as replaying older examples or constraining updates remain proposals until tested under the same retained suite. The forgetting paper motivates methods, not a guarantee of preservation for an LLM agent. Any remedy should be evaluated for new-task gains and old-task retention together.

The practical artifact is a retention matrix with before-and-after counts and the affected instruction families. We would accept a claim of improved compliance only in the families and conditions measured. A higher overall score alone cannot establish that an agent still respects the owner's earlier boundaries.

Sources

Related field notes