Attach user feedback to the trace it describes
By DX Research Group · · Data and learning flywheels
A proposed evidence bundle preserves the decision, configuration and timing behind each feedback label.
A feedback sentence needs a referent. “The agent ignored me” can describe an old instruction, a rejected proposal, a confusing explanation or an actual action outside the intended mandate. We propose attaching feedback to an immutable decision record before using it to evaluate or improve an agent.
DXAP's activity guide gives owners a path from a selected turn to action details and account outcomes. The chat guide distinguishes conversation from confirmed instructions and approved settings. Those distinctions determine which evidence a feedback record must preserve.
An illustrative late report
Suppose a turn begins at 09:00 under strategy version A. At 09:03 the owner approves version B, adding a new condition. At 09:05 the owner flags the earlier turn as ignoring that condition. Joining the report to the latest strategy would manufacture a compliance failure. The case needs version A as the instruction available to the turn, version B as the later owner change, and the relevant execution-time configuration if an action continued after approval.
Our proposed bundle would retain the feedback text and submission time alongside a stable turn identifier. It would reference the rendered input, instruction version and applicable policy state through controlled internal storage. The owner could select the specific explanation or action being criticized. Evidence availability would be recorded explicitly: a missing input snapshot leaves the interpretation unresolved even if the complaint is persuasive.
Three judgments can then coexist. The owner wanted a different behavior; the earlier model may have followed version A; and a later action may still require checking against newly effective execution limits. Preserving this sequence prevents a single negative rating from becoming three unsupported labels.
The controls paper companion describes historical invocation-level linkage across mandate, prompt, action, validation and outcome. We would extend that lineage with feedback, recording who supplied the interpretation and which evidence was available to that reviewer. A source-linked interpretation remains an interpretation until the underlying behavior is checked.
Keep corrections append-only
Imagine the first reviewer marks the case “instruction ignored,” then finds that version B was approved later. The corrected label should supersede the earlier judgment while preserving its reason and timestamp. Silent replacement would make a later evaluation impossible to reconstruct. A dataset export should declare whether it contains original labels, final adjudications or both.
Privacy belongs in this design at collection time. Internal references can support authorized reconstruction without exporting account identifiers or complete conversations. Shared research fixtures should contain only the context needed to reproduce the failure, with permission appropriate to that use. A public article can explain the mechanism using a synthetic case.
We would audit a sample of feedback records for successful reconstruction, version correctness and reviewer agreement. Duplicate reports attached to one turn remain multiple perspectives on one event. They can improve interpretation without multiplying the event count.
This proposal makes user feedback useful as evidence about the decision system. It supplies the provenance required for a later harness experiment or training dataset, while leaving the choice of intervention open. The immediate improvement to test is whether reviewers diagnose the same case more consistently when they receive the complete, time-correct bundle.