Separate Model Uncertainty from Market-Data Uncertainty
By DX Research Group · · Learning theories
More model agreement cannot repair a stale or ambiguous observation.
Uncertainty about a market observation and uncertainty about a model's interpretation require different interventions. We would preserve both in an evaluation rather than treating a single confidence score as a complete diagnosis.
Kendall and Gal distinguish uncertainty inherent in observations from uncertainty in the model in computer-vision tasks. Applying that distinction to trading needs care: feed outages, source conflicts, and genuinely unpredictable future returns are different forms of uncertainty. Our proposed fixture separates them operationally. Trace feedback records the inputs, and harness-transfer tests support controlled substitutions.
A variance calculation with a clear target
For an illustrative forecast Y, suppose two equally weighted model hypotheses predict means of 2 and 6 units. Each predicts conditional outcome variance 9. The mean forecast is 4. The average conditional variance is 9; variance of the two means is ((2 minus 4) squared plus (6 minus 4) squared)/2, or 4. Total predictive variance is 13 under this mixture.
The calculation follows the law of total variance. Calling the 4-unit component model uncertainty depends on whether the hypotheses represent meaningful uncertainty about the model. Two arbitrary prompts on the same model may measure prompt sensitivity instead. Similarly, the 9-unit component bundles whatever conditional noise the forecast distribution models; it cannot identify a stale feed by itself.
A source-quality flag adds information outside that decomposition. If both models receive the same incorrectly timestamped quote, they can agree precisely on a bad premise. Their low disagreement should coexist with a data-validity failure.
A two-axis fixture
We would prepare four conditions: clean data with a familiar task, clean data with an unfamiliar task, corrupted data with a familiar task, and corrupted data with an unfamiliar task. Corruption is controlled and visible in the fixture, such as a quote timestamp moved outside its permitted age. Task unfamiliarity is defined from the registered training population, rather than assigned after observing an error.
Measure forecast distribution quality, detection of invalid inputs, and the resulting action separately. A model that admits uncertainty while submitting the same overconfident action needs its decision policy examined. A model that abstains correctly on bad data may still have weak predictive discrimination on valid data.
The fixture also tests recovery: replace the corrupted quote with a valid saved observation and check whether the data warning resolves. Changing only the model should leave the source-invalid status intact. The outcome is a diagnostic map from uncertainty source to response, with economic performance evaluated in its own population. It gives us a specific next intervention rather than a general instruction to make the model more confident.