Learn the owner’s preferences without learning yesterday’s price move
By DX Research Group · · Data and learning flywheels
A proposed preference dataset turns explicit owner choices into response targets while withholding future market outcomes.
An owner can teach an agent how to communicate without teaching it which asset will rise. That distinction gives user feedback a useful training destination. We propose a preference dataset whose targets come from explicit judgments about responses, with future price movement withheld from the labeling task. The objective is better owner alignment within the trading harness.
DXAP’s chat guide describes conversations grounded in the selected agent’s strategy, account and recent activity. It also distinguishes conversation from trade execution and approval. Those documented boundaries provide a practical place to collect preference examples, subject to consent and review. They establish interaction mechanics; the training loop described here is a proposed extension.
One correction, two possible lessons
Consider a fictional owner who says, “Show the missing entry condition first. Put the market recap afterward.” The preceding response explained a waiting decision accurately, but buried that condition in its fourth paragraph. The owner’s correction supplies a response-order preference. A curator can create a target that states the unmet condition first and preserves the same facts and action.
Suppose the asset rises afterward. Adding “waiting was wrong” to this training row would mix a later market outcome into a correction about presentation. The row should contain the question, the point-in-time evidence, the original response and the approved rewrite. Its label names the requested presentation change. Subsequent prices belong in a separate forecast or economic evaluation with its own horizon.
A useful paired fixture keeps the market packet identical and changes the owner preference. Owner A requests the missing condition first; owner B requests a compact exposure summary first. The expected response differs in ordering while both retain the same decision and policy facts. Success means the model adapts to the declared preference rather than treating one owner’s style as a universal rule.
For an illustrative audit, imagine twenty preference examples, five of which require a response to retain a material uncertainty sentence. A candidate that follows the requested order on eighteen examples but deletes uncertainty in three of those five has learned an incomplete lesson. Report order adherence and uncertainty preservation separately. The owner’s desire for brevity supplies no permission to erase a decision-relevant fact.
Give preferences a scope
The curation record should say whether a correction applies to one answer, one agent or a reusable owner preference. An isolated frustrated comment has weaker scope than an explicit persistent choice. Conflicting preferences need their conditions restored: a long explanation may suit a strategy review while a short answer suits a routine activity question.
We would hold out owners and question families when evaluating a shared update. An owner-specific adaptation can use a separate within-owner test, with examples collected before the evaluation window. This distinguishes remembering an individual’s instruction from generalizing a response skill to new people.
The controls paper companion describes how earlier agent decisions could acquire authority they were never given. A preference pipeline should avoid a parallel mistake: turning yesterday’s favorable outcome into today’s owner instruction. Curated labels make the intended lesson explicit. The measurable advance would be fewer owner-requested rewrites at equal factual coverage, with forecasting and simulated economics assessed independently.