Native probabilities versus verbal confidence in trading agents

By DX Research Group · · Frontier research

A proposed calibration study separates token probabilities, spoken confidence, market forecasts and decisions under one saved evaluation set.

A trading agent can sound certain while assigning weak probability to its own answer. We want to measure that discrepancy before turning confidence into position size. This is a PROPOSED experiment, with no reported outcome or commitment that DXAP will ship a probability head. Its object is a defined forecast, such as whether a specified asset closes above its current reference price over the next hour, rather than the general persuasiveness of a trade explanation.

Define the event before asking the model

We would save the market snapshot, forecast horizon and settlement rule for every question. The label must use a price series available under a documented timestamp convention. Questions should include quiet periods, volatile periods and occasions when the agent would choose to observe. Dropping observations because the agent declined a trade would turn a forecasting study into a selected trade study.

For one arm, the model emits a verbal probability in a fixed response field. For the other, an instrumented route returns native probabilities for a constrained candidate answer. We would check tokenization and preserve the full candidate probability mass. If an answer occupies multiple tokens, the probability of its first token alone is an inadequate score for that answer. A provider route that cannot expose the required probabilities belongs outside that arm.

A confidence number has a specific meaning

An illustrative fixture contains four forecasts: 0.55, 0.60, 0.70 and 0.80, with outcomes 0, 1, 0 and 1. Their mean squared probability error is 0.248125. That is a calculation on invented values, included to make the scoring rule inspectable. It establishes neither calibration nor market skill. On saved data we would compare Brier score and log loss against a constant base-rate forecast, then inspect reliability across probability ranges.

Calibration fits should use earlier windows and be evaluated on later untouched windows. We would publish the number of events and the dependence between events, because overlapping one-hour targets can make a large row count look more informative than it is. Market-day clustered comparisons help preserve that dependence when reporting uncertainty.

Forecasting and trading need separate verdicts

Our continuous record found no directional edge in the historical fleets. A better probability score would therefore require direct evidence rather than a confident explanation. We would hold entry thresholds, sizing and exit rules fixed when evaluating simulated decisions, reporting forecasting improvement separately from net results after explicit costs. A calibrated forecaster can still trade badly if its decisions are expensive.

The earlier controls paper explains why the surrounding harness matters. DXAP currently describes selectable models inside a common harness and recorded decisions. That architecture gives a concrete reason to compare confidence channels under identical inputs. The useful evolution would be an inspectable forecast record whose meaning survives a model change, supported by held-out results before confidence influences capital.

Sources

Related field notes