Teach mandate clarification before rewarding actions
By DX Research Group · · Data and learning flywheels
A proposed training curriculum uses owner-confirmed clarifications to teach the objective an action must serve.
An action reward is difficult to interpret when the owner’s request remains ambiguous. We propose teaching mandate clarification first, using owner-confirmed interpretations as supervised examples. The resulting objective is to recognize when an instruction is ready for action and when one missing detail changes its meaning.
DXAP’s configuration reference distinguishes strategy text from configured trading limits. The chat guide describes reviewed settings proposals and confirmed persistent instructions. Together they supply a documented interaction boundary around which a clarification dataset could be designed. Current documentation establishes those product mechanics, while the curriculum here remains a research proposal.
Imagine a fictional request: “Keep each entry small, around five hundred.” The missing unit matters. A curated example should reward asking whether the owner means $500 notional and whether that is a maximum or an approximate target. It should preserve the owner’s eventual answer, such as “Maximum $500 notional per opening trade,” as the confirmed interpretation.
The training target should describe the proposed interpretation and the required approval step. It should avoid pretending that clarification alone saved a policy setting. A second example starts with an already explicit request: “Propose a maximum entry notional of $500.” Repeating the same unit question here wastes the owner’s attention. This contrast teaches selective clarification instead of universal hesitation.
An ordered curriculum
We would first train on ambiguity detection and focused questions. Next, train on the mapping from a confirmed interpretation to a proposed instruction or supported settings change. Only then consider action objectives evaluated against the confirmed mandate and actual saved policies. Each stage receives its own target, so an attractive action cannot compensate for a mistaken reading of the owner’s request.
For an illustrative exercise, twelve requests contain genuine ambiguity and twelve are sufficiently explicit. A candidate that asks questions for all twenty-four catches every ambiguity but creates twelve unnecessary interruptions. Report missing necessary questions and unnecessary questions as separate counts. The curriculum should improve the tradeoff at a declared interaction budget, rather than celebrate one rate alone.
User contribution supplies the missing semantic information. Curators must retain the initial wording, the clarifying question and the confirmed answer as one linked example. Training only on the final mandate removes the ambiguity that made the question necessary. Training on a support agent’s guess without owner confirmation gives the label more authority than its source warrants.
Keep authority outside the curriculum
The policy reference describes execution checks on proposed orders. A model update should preserve those checks. A learned ability to interpret “small” supplies a candidate meaning; authorization still comes from the owner’s confirmed instruction and saved configuration.
Evaluation would include new phrasings, conflicting context and explicit units across held-out mandate families. Follow the clarification through to the proposal and verify that unchanged fields remain unchanged. Economic evaluation can then ask what compliant decisions do under a fixed execution policy, using a separate outcome dataset.
This sequence makes the learning objective legible. Users contribute precise intended meanings; supervised examples teach the model how to obtain and preserve them. The proposed first advance is fewer mistaken interpretations with fewer unnecessary questions, before any claim about rewarded trading actions.