Reward Normalization Across Market Days Changes the Lesson
By DX Research Group · · Learning theories
Per-day standardization can erase an economically important difference between quiet and volatile sessions.
Normalizing rewards by market day changes which errors receive training pressure. It can make optimization numerically easier while replacing an economic objective with relative standing inside each day. We would write down both objectives before choosing the transform.
Implementation choices in reinforcement learning can materially affect reported agent performance. This proposed trading-specific calculation concerns one such choice. Our continuous-record companion keeps historical economics attached to their measurement setting, while trace feedback distinguishes local diagnosis from downstream value.
Equal standardized rewards can hide unequal costs
Consider two illustrative days with two observations each. Day A has rewards 1 and 3 units. Its mean is 2 and its population standard deviation is 1. Day B has rewards minus 10 and 30 units. Its mean is 10 and standard deviation is 20. Standardizing each day produces minus 1 and plus 1 for both days.
A learner using these transformed targets sees equally sized positive and negative deviations. The raw rewards distinguish a 2-unit spread on A from a 40-unit spread on B. Neither scale is universally correct. The choice depends on whether the objective values relative quality within sessions or absolute consequences under a fixed economic unit.
Normalizing by realized daily volatility introduces another issue. A completed day's scale includes events after an early decision. That may be acceptable for an explicitly retrospective training target, but it cannot be presented as a decision-time feature. The transform needs a declared information boundary.
We would compare a fixed, training-only global scale with a trailing scale computed using previously available days. Keep raw outcomes in the dataset so the original economic meaning remains recoverable. Fit any clipping threshold on the training period, then freeze it for validation. A near-zero scale requires an explicit handling rule because a tiny denominator can amplify ordinary noise.
Check gradient emphasis and economic evaluation separately
The proposed diagnostic reports how much total sample weight each market day receives after transformation. A day with many turns may still dominate even after its rewards are standardized. If the intended unit is a day, average within day before averaging across days; if the intended unit is a decision, retain the decision-weighted summary and explain that choice.
Evaluate a candidate with raw net outcomes under the same fixed decision policy, alongside the transformed optimization score. Report mandate compliance and exposure as their own measures. A higher standardized reward can coexist with worse raw outcomes if the transform changes which situations the model prioritizes.
Before any training, we can inspect the transformation on saved targets and display the examples whose ranking changes most. That audit makes the proposed objective concrete. The decision to use per-day normalization then has a visible consequence rather than appearing as an innocuous preprocessing default.