An owner correction needs a curator decision
By DX Research Group · · Data and learning flywheels
A proposed rubric turns precise owner feedback into bounded labels instead of assumed ground truth.
An owner can be authoritative about their intended instruction while still being mistaken about an observed trade or a forecast. A contribution curator needs to separate those roles before turning feedback into a target answer. We propose accepting corrections by claim type, with a reproducible reason attached to each accepted label.
DXAP's policy and execution documentation distinguishes a successful policy check from a venue fill and from profitability. That distinction should carry into the correction queue. “The agent ignored my cap,” “the venue filled too much” and “the market moved against me” require different evidence, even when they arrive in the same message.
Review the claim, then the evidence
Our proposed rubric asks four questions: what claim is being corrected, what source establishes the expected behavior, whether the saved input is sufficient to reproduce it, and what permission covers the resulting example. The curator can accept one claim while leaving the others unresolved. A whole-report thumbs-up would erase that useful granularity.
Consider an illustrative owner instruction: “Use at most 3% of available equity for a new entry.” The saved snapshot shows $10,000 of available equity and a proposed $400 entry. The expected cap is $300. If the instruction was effective at the proposal time and the field means entry notional, the contribution supports an instruction-compliance label. The arithmetic is reproducible and the authority comes from the authenticated instruction.
Now add the owner's sentence: “A $300 trade would have won.” That is a separate economic claim. Establishing it requires a declared counterfactual entry, exit policy and cost model. The curator can accept the cap violation while marking the profit claim unsupported by the supplied artifacts. A useful owner correction survives without turning hindsight into a prediction target.
If the report instead says “I meant 3% margin,” the proposed label concerns ambiguous instruction interpretation. It may motivate a clarification fixture, rather than relabeling the original model output as an obvious error. Retrospective intent should be recorded as later clarification with its own timestamp. The evaluation can then ask whether the harness should have requested clarification before proceeding.
Make disagreements informative
We would have two reviewers independently assess a sample, recording disagreement by claim type. A disagreement about instruction authority calls for better provenance. A disagreement about arithmetic calls for a deterministic calculation. A disagreement about forecast quality belongs in a prediction evaluation with observed outcomes. Agreement percentages alone conceal these different repairs.
The controls paper companion motivates locating failure along the mandate-to-settlement path. The proposed rubric adds owner feedback without making owner confidence a universal truth label. Its first success measure would be reproducible acceptance decisions and lower unresolved ambiguity in the reviewed population. Whether accepted examples improve a later model or harness release remains a held-out evaluation question.