Describe a learning roadmap through assessed capabilities
By DX Research Group · · Data and learning flywheels
A stage-specific acceptance table makes future adaptation concrete without treating a roadmap as a shipped learning system.
A learning roadmap becomes useful when each stage names a capability that can be assessed. “The agent learns from use” leaves too much unspecified: whose experience, which component changes and how an owner can recognize a regression? We propose describing DXAP's next learning advances through separate acceptance records for remembering, diagnosing, adapting and selecting releases.
There is already a concrete starting point. DXAP's October 3 release notes describe agent-scoped chat memory when enabled, including requests to save, correct or forget a detail. They also describe strategy reviews that distinguish evidence under earlier configurations from behavior under updated settings. These shipped features give owners continuity and a clearer basis for steering. They establish a place to start testing richer adaptation; online model-weight updates remain a proposed research direction here.
Four assessed capabilities
Our proposed acceptance table gives each capability a different receipt:
| Capability | Change being assessed | Receipt required |
|---|---|---|
| Remember | Durable context for one agent | Correct save, correction and forgetting on new conversations |
| Diagnose | Classification of a reported failure | Correct cause on resolved incidents whose answers were withheld |
| Adapt | A revised retrieval, harness or candidate model artifact | Improvement on unseen cases with stable owner-control behavior |
| Select | A process choosing whether a candidate advances | Correct rejection of deliberately harmful candidate releases |
These stages describe assessment scope, rather than a promised shipping order. A platform can improve diagnosis before changing model weights. A retrieval repair may be sufficient for a memory issue. Each result should name the artifact that changed so readers can tell what produced it.
Consider an illustrative owner whose preference changes from concise explanations to detailed explanations. A memory test checks whether the next conversation uses the new preference and forgets the old one when asked. A diagnosis test asks whether the system correctly identifies an obsolete retrieved preference as the cause of a later mismatch. An adaptation test evaluates a revised retrieval rule on unfamiliar preference changes. A selection test offers a candidate that fixes verbosity while silently applying a strategy edit. The appropriate selection is rejection, even if its conversational score improves.
The execution documentation supplies the stable contract: configured checks apply before proposed orders reach the venue. Learning acceptance should include those controls as independently assessed behavior. Conversational success has its own measure; forecast quality needs time-valid labels; economics requires costs and execution outcomes. A roadmap can carry all three without combining them into a single progress percentage.
Our controls research shows why staged receipts matter. Bounded harness interventions changed specific trace behavior, which gives the next experiment a mechanism to test. A useful milestone would say that a candidate localized an unfamiliar failure family correctly and preserved authorization on adjacent cases. Its successor would show that the resulting repair generalized. That is a roadmap owners can inspect as the platform evolves.