Evaluate feedback in a shadow path before changing the agent
By DX Research Group · · Data and learning flywheels
A proposed shadow study compares feedback-aware candidates while preserving the live authority path.
User feedback can be evaluated without immediately changing the agent that controls an account. We propose a shadow path that receives an authorized copy of selected decision inputs, produces an explanation or intended action, and records the comparison. It has no execution credential and no route to modify the owner's settings.
The distinction fits the published DXAP decision path: an eligible turn evaluates context, proposes an action or waits, and execution checks authorization and configured policies. The chat guide separately requires owner review for proposed settings and persistent instructions. A feedback-aware research candidate should preserve those relationships throughout evaluation.
Two outputs from one decision opportunity
Take an illustrative report that the agent's explanations omit an important uncertainty in a data source. The proposed shadow candidate receives the same saved inputs as the incumbent plus an authorized, curated feedback example about stating missing coverage. Both produce their intended action and explanation. The live agent continues through its ordinary path; the candidate output is retained as research evidence.
Suppose a saved input contains a quote with a known timestamp but missing depth. The incumbent proposes an entry and cites the price. The candidate also proposes an entry but states that depth coverage is unavailable. An explanation improvement may have occurred. To claim a decision improvement, evaluators must also inspect whether that uncertainty changes the proposed action appropriately under the unchanged mandate. To claim an economic improvement, they need a separate outcome study with declared execution assumptions.
A candidate that changes the maximum entry size in its own output has created a different question. The evaluator should record the attempted authority change even if a downstream policy check would reject the resulting order. Research feedback supplies information about an error; it supplies no permission to rewrite limits.
Measure disagreement where it occurs
We would register three paired measures: factual completeness of the explanation, agreement with the saved mandate, and validity of the intended action under the common policy. Review a blinded sample of disagreements and keep missing candidate outputs in the coverage denominator. Any simulated economic comparison should use one declared fill model and common costs, with its assumptions attached to the result.
Shadow outputs also need an input boundary. A candidate evaluated after the live outcome arrives could benefit from later fills or prices. Its input copy must freeze at the same decision cutoff. Feedback used to build the candidate belongs in the development record; feedback first observed after that cutoff cannot be passed into that historical decision as though it were available then.
Our proposed operational check is straightforward: the shadow identity cannot submit an order, approve an instruction, or save a settings mutation. Attempted calls to those routes must fail and be recorded during rehearsal. Evaluating that capability boundary is part of the study, alongside grading the model's text.
The benefit is a measured account of how a feedback-aware candidate differs on real decision opportunities. A passing shadow study would support a later release review within its tested scope. Actual deployment would remain a separate versioned decision under the same owner control contract.