Building a Time-Valid Base-Rate Baseline for Agent Forecasts

By DX Research Group · · Forecast evaluation

A fixed reference probability prevents an impressive accuracy number from becoming a misleading skill claim.

We define our baseline before opening evaluation labels so the timing remains inspectable.

Define a reference that could have existed

An LLM forecast needs a comparison that is available at forecast time. A constant event-frequency reference is often informative because many event definitions are imbalanced. If positive events are rare, high classification accuracy can arise from predicting the negative class every time.

The scikit-learn DummyClassifier documentation specifies that its prior strategy produces the empirical training class distribution. The important research distinction is training. Estimating a reference probability from the evaluation outcomes gives it information that was unavailable when the agent issued its forecasts.

Calculate the reference under a changing frequency

Suppose an earlier training window contains 200 resolved questions with 40 positives. Its observed frequency is 0.20. Freeze that probability before evaluating 100 later questions. Suppose the later window contains 30 positives and 70 negatives. These are illustrative counts.

The frozen reference's Brier score is (30 times 0.80 squared plus 70 times 0.20 squared) divided by 100. This equals (19.20 plus 2.80) divided by 100, or 0.22.

A constant 0.30 forecast fitted to the evaluation labels would score 0.21. That better number describes an in-sample frequency estimate. It is useful as a descriptive diagnostic but it is not a deployable baseline for the same evaluation window. Label it separately if shown.

Meanwhile, always classifying the event as negative yields 70% accuracy. A model with 68% accuracy could still provide better probabilities than the 0.20 reference, while a model with 75% accuracy could provide worse probabilities. The accuracy figure cannot settle the probability comparison.

Choose updating rules before labels arrive

A rolling reference can adapt to changing frequency, but specify its window, minimum count, eligibility filters, and update timing. If a question resolves late, it cannot enter the reference until its label is available. A row's start time is not its resolution time.

Consider a daily update using the most recent 200 resolved questions. Save the actual count and probability used each day. If unresolved questions are excluded, record their number and age. Otherwise an apparently changing base rate may partly reflect which types of questions resolve quickly.

Use the same target and eligible question set as the model comparison. Separate instruments or horizons only when the reference rule specifies them in advance and there is sufficient support. A different reference for every tiny subgroup can overfit historical variation and become unstable.

Interpret what beating the reference establishes

Compare paired losses on identical saved questions. A lower aggregate Brier or log loss indicates better probability scoring on that sample relative to that specific reference. It does not show that every subgroup improved, that the benefit survives another period, or that agent actions become economically favorable.

For agentic trading, a base-rate reference is valuable precisely because it does little. It isolates what the model adds beyond historical event frequency. Keep execution costs, mandate compliance, and action outcomes in separate evaluation fields. That separation makes a forecast comparison interpretable without promoting it into a product performance claim.

The practical deliverable is a baseline specification with a dated training boundary, label availability rule, update schedule, frozen predictions, and paired score table. A reference reconstructed after viewing the test outcomes cannot serve the same purpose.

Connect the metric to the research record

Our agent evaluation framework separates the forecast comparison from the action and execution layers. The continuous production record shows why reporting an unfavorable result is part of a useful research account.

Sources

Related field notes