Assigning Delayed Credit to Agent Actions
By DX Research Group · · Learning theories
A later portfolio outcome needs a declared path back to the decisions that could have affected it.
Delayed outcomes should be attributed through an explicit action sequence and time horizon. Giving every earlier turn the final portfolio reward can teach a tool call or no-trade decision from an event it could barely influence. We propose a dependency-aware credit audit before choosing a learning target.
Safe and Efficient Off-Policy Reinforcement Learning addresses learning from sequential experience. Its general sequential setting motivates this question, while the attribution protocol here is our proposed trading fixture. Our trace-feedback framework supplies linked stages; the continuous record retains its historical population and execution limits.
Three turns, one eventual outcome
In an illustrative sequence, turn 1 opens a position, turn 2 retrieves fresh funding information, and turn 3 closes half the position. A later mark-to-market gain of 12 units belongs to the evolving exposure across those intervals. It does not by itself establish that retrieving the funding information created 12 units of value.
Suppose exposure is 2 contracts for the first interval and 1 for the second. Price moves plus 4 units per contract during the first interval and plus 4 during the second. Ignoring costs, interval attribution yields 8 plus 4, totaling 12. This accounting allocates exposure consequences; it still does not identify the counterfactual contribution of the information request.
For that tool call, we need a separate intervention: replay the later decision with and without the saved tool result, holding the market path and eligible information cutoff fixed. If the action remains identical, the tool may improve explanation or verification while producing zero measured action-mediated effect in that fixture.
A terminal reward assigned to the entire sequence can also cross mandate changes. An owner might authorize closure halfway through the episode. The episode boundary should record which decisions had authority over each exposure interval and which later events were externally imposed.
Choose the credit question explicitly
We would retain three views: settlement accounting by exposure interval, decision-quality labels using information available at each turn, and proposed counterfactual action effects under a declared simulator. The views answer different questions. Combining them into one scalar should require an explicit objective and preserved component values.
The audit includes delayed fills, canceled orders, and later corrections. An accepted proposal may fail to create exposure; a stale order may fill after a subsequent decision. Attribution follows the economic event's action lineage instead of the nearest timestamp.
A useful first artifact is a small episode diagram rendered as a table of turns, exposure changes, and outcome windows. Each reward target points to its supported dependency path. Where the market simulator cannot support an alternative action, the counterfactual stays unknown. This keeps delayed learning constructive while preserving the distinction between cash-flow accounting and causal credit.