Recording the Trials Behind a Winning Agent Forecast Score
By DX Research Group · · Forecast evaluation
A simple chance calculation explains why a selected winner needs an untouched evaluation.
We keep our candidate search history attached to the eventual development winner.
The winner inherits the search
Comparing native LLM prompts, models, probability maps, and policy thresholds creates a selection process. The best observed score is the result of both candidate quality and the search over noisy measurements. A report that presents only the winning candidate leaves reviewers unable to distinguish those contributions.
Cawley and Talbot's model-selection paper studies overfitting of selection criteria and subsequent evaluation bias. The limited claim used here is that model selection itself can overfit a finite evaluation sample. The numerical illustration below is independent arithmetic under stated assumptions.
A small chance calculation
Suppose a hypothetical test procedure has a 5% false-positive probability for each truly null candidate. Under the simplifying assumption that candidate tests are independent, the probability that none of 20 candidates produces a false positive is 0.95 raised to the twentieth power, approximately 0.3585. The probability of at least one is therefore approximately 0.6415, or 64.2%.
For 100 independent tests, the corresponding probability is 1 minus 0.95 raised to the hundredth power, approximately 0.9941, or 99.4%. These calculations do not say a real hundred-model leaderboard has that error rate. Real candidates are correlated, tests may use different criteria, and selecting the lowest score is not identical to selecting a significant test. The example isolates why repeated opportunities matter.
Trying five prompts, four temperatures, and five calibration settings creates 100 combinations if every combination is examined. Renaming the process “iteration” does not remove those selection opportunities. Nor does using the same nominal base model make the variants statistically identical.
Keep a ledger with failures
Record each candidate's model version, prompt version, input population, probability transformation, selection metric, evaluation period, and reason for inclusion. Include candidates that timed out or produced invalid outputs. They affect coverage and can influence the final choice even when no valid score is produced.
Separate development data used to rank variants from a final evaluation period. Freeze the complete selected procedure before opening that period. If a final result prompts another prompt revision, the period has become development evidence for the revision. A new untouched comparison is needed for the revised procedure's prospective claim.
The ledger should also record informal inspection. Looking at individual test failures and rewriting the prompt uses test information even if the aggregate metric was never optimized automatically. Human selection can overfit as readily as a grid search.
State the actual decision supported
A development winner can justify selecting one candidate for a further test. Reliable superiority requires that separate test. A held-out improvement can support a scoped comparison on that population, but uncertainty, temporal change, and execution assumptions still limit the conclusion.
For frontier agentic trading research, distinguish the forecast selection procedure from the action policy. If the same outcomes choose both a probability map and an entry threshold, both choices consume the sample. Report the full procedure and a fixed baseline on identical questions.
A practical publication can show the number of variants tested, their score distribution, the frozen selection rule, and the final comparison. This gives readers a meaningful denominator for a winning score without inventing a correction factor or treating a single favorable run as proof of product advantage.
Connect the metric to the research record
Our agent evaluation framework separates the forecast comparison from the action and execution layers. The continuous production record shows why reporting an unfavorable result is part of a useful research account.