Describe a learning roadmap through assessed capabilities

By DX Research Group · · Data and learning flywheels

A stage-specific acceptance table makes future adaptation concrete without treating a roadmap as a shipped learning system.

A learning roadmap becomes useful when each stage names a capability that can be assessed. “The agent learns from use” leaves too much unspecified: whose experience, which component changes and how an owner can recognize a regression? We propose describing DXAP's next learning advances through separate acceptance records for remembering, diagnosing, adapting and selecting releases.

There is already a concrete starting point. DXAP's October 3 release notes describe agent-scoped chat memory when enabled, including requests to save, correct or forget a detail. They also describe strategy reviews that distinguish evidence under earlier configurations from behavior under updated settings. These shipped features give owners continuity and a clearer basis for steering. They establish a place to start testing richer adaptation; online model-weight updates remain a proposed research direction here.

Four assessed capabilities

Our proposed acceptance table gives each capability a different receipt:

CapabilityChange being assessedReceipt required
RememberDurable context for one agentCorrect save, correction and forgetting on new conversations
DiagnoseClassification of a reported failureCorrect cause on resolved incidents whose answers were withheld
AdaptA revised retrieval, harness or candidate model artifactImprovement on unseen cases with stable owner-control behavior
SelectA process choosing whether a candidate advancesCorrect rejection of deliberately harmful candidate releases

These stages describe assessment scope, rather than a promised shipping order. A platform can improve diagnosis before changing model weights. A retrieval repair may be sufficient for a memory issue. Each result should name the artifact that changed so readers can tell what produced it.

Consider an illustrative owner whose preference changes from concise explanations to detailed explanations. A memory test checks whether the next conversation uses the new preference and forgets the old one when asked. A diagnosis test asks whether the system correctly identifies an obsolete retrieved preference as the cause of a later mismatch. An adaptation test evaluates a revised retrieval rule on unfamiliar preference changes. A selection test offers a candidate that fixes verbosity while silently applying a strategy edit. The appropriate selection is rejection, even if its conversational score improves.

The execution documentation supplies the stable contract: configured checks apply before proposed orders reach the venue. Learning acceptance should include those controls as independently assessed behavior. Conversational success has its own measure; forecast quality needs time-valid labels; economics requires costs and execution outcomes. A roadmap can carry all three without combining them into a single progress percentage.

Our controls research shows why staged receipts matter. Bounded harness interventions changed specific trace behavior, which gives the next experiment a mechanism to test. A useful milestone would say that a candidate localized an unfamiliar failure family correctly and preserved authorization on adjacent cases. Its successor would show that the resulting repair generalized. That is a roadmap owners can inspect as the platform evolves.

Sources

Related field notes