What changes when an agent learns?

By DX Research Group · · Data and learning flywheels

A change-location test separates saved context, harness revisions, and proposed model training.

“Self-learning” becomes useful when we can identify the object that changed. A conversation can update saved context. A developer can revise the harness around the model. A training process can update model parameters. These mechanisms have different release obligations, and each requires its own evidence.

DXAP already documents one concrete mechanism. Its October 3 release notes describe agent-scoped chat memory: when enabled, preferences, goals, and user-adopted theses can carry into new conversations with that agent. Owners can ask it to save, correct, or forget a detail, and background processing can capture durable context. The documentation keeps this memory separate from trading notes and from strategy or settings changes. That is a shipped context mechanism. The inspected record establishes no current automatic weight-training loop.

Locate the changed object

We propose a change-location test for future learning claims. Start with a saved question, a fixed account snapshot, and a versioned configuration. Run the same question before and after the claimed improvement. Record which persistent object differs: memory entries, harness instructions or tools, model parameters, or several together. A provider model identifier alone may be insufficient to establish a weight revision; the release needs an identifiable model artifact or provider version boundary.

An illustrative owner says, “Remember that I prefer explanations in USDC.” A later conversation uses USDC without another reminder. Clear the relevant memory in a test copy and repeat the question. If the adaptation disappears while the model and harness remain fixed, the evidence points to memory. This test measures retrieval and use of a preference. It says nothing about market forecasting.

Now consider a corrected explanation of a stale quote. A revised harness could require the quote's observation time beside its price. If the fix persists across fresh test agents that share the new harness and have no contributed memory, its scope is a release-level change. A proposed weight update would need a separate comparison with the same harness and memory inputs on both sides.

Attach a measurement to each mode

For memory, measure whether a permitted detail is stored, retrieved for the right agent, corrected, and forgotten on request. For a harness release, measure the targeted behavior on independent cases and record its version. For a proposed trained model, compare candidate and incumbent on untouched questions with common tools and inference budgets. Keep instruction compliance, forecast scoring, and simulated economics in separate outputs.

The controls paper companion provides historical evidence that harness revisions changed specific trace behaviors. Its combined rule-fabrication intervention and fee-placement experiment concern the recorded behavior of the surrounding system. They supply a reason to inspect the changed object carefully rather than assigning every improvement to the model.

A useful release description would therefore read: “Agent-scoped preference recall improved in this tested memory path,” or “This harness revision reduced this identified error.” Our proposed test makes those statements checkable. DXAP's distinctive opportunity is the connection between persistent agent context and inspectable decisions: adaptation can leave a record that identifies what changed and which owner controls remained in force.

Sources

Related field notes