Probability Calibration and Trading Policy: Test Them Separately
By DX Research Group · · Learning theories
A worked example shows how a calibrated forecast can produce different trading actions under different cost and risk policies.
A probability forecast and a trading action answer different questions. The forecast estimates an event. The policy decides whether a position fits the expected payoff, execution cost, and mandate. We can improve one while the other stays weak, so a learning experiment needs two separate comparisons.
Consider an illustrative agent forecasting whether a market price will rise over the next hour. It outputs 0.60. A calibrator maps that score to 0.54 based on held-out historical observations. A trading policy then chooses size or observation from the calibrated score, expected move size, costs, and current exposure. The change in probability has a precise meaning; its trading consequence depends on the policy.
A Forecast Alone Leaves Payoffs Open
Suppose the favorable move pays $10 and the unfavorable move loses $10, before costs. At a 0.54 probability, the illustrative expected gross payoff is $0.80: 0.54 times $10 minus 0.46 times $10. If the estimated round-trip cost is $1, the expected net payoff becomes minus $0.20.
Now change the favorable payoff to $20 while leaving the loss at $10. The same probability gives an expected gross payoff of $6.20, before that $1 cost. A single confidence threshold would treat these situations identically despite different payoff structures.
The example assumes known payoffs for clarity. In real trading, move size and execution costs are estimates with their own uncertainty. The experiment should preserve those estimates as they were available at decision time.
Freeze One Component at a Time
Our proposed first comparison holds the action policy fixed and replaces only the calibrator. Evaluate probabilities on a held-out temporal split with a declared event definition and proper probability score. Inspect calibration by score band with sample counts, because a smooth chart can conceal sparse observations.
The second comparison holds the forecast and calibrator fixed while changing the action policy. Evaluate exposure, turnover, cost burden, and simulated economic outcomes under identical accounting. This isolates policy behavior from predictive quality. A combined update can be useful later, once the individual contributions are understood.
Select the probability horizon before viewing results. A forecast of a one-hour sign should be scored against that event even if the position exits sooner. If the policy needs a different event, formulate and test that forecast explicitly rather than retroactively changing the label.
What This Means for Agent Evolution
Our evaluation method asks for the complete decision system and its measurement boundary. The harness-transfer article supplies the component-level comparison structure needed here. This note contributes a specific separation between probability mapping and action selection.
DXAP's public loop places policy checks outside the model and records the resulting turn. That separation is useful for inspecting whether an action stayed within limits. It says nothing by itself about forecast calibration, which requires a separately labeled dataset and scoring experiment.
We would report improved calibration as a forecast result. We would report better policy outcomes as a decision result under stated assumptions. Product evolution becomes more informative when those two results retain their own evidence.