Macro and Micro Averaging Across Trading Agents
By DX Research Group · · Forecast evaluation
Choose whether the research question concerns a typical agent or a typical forecast.
Micro averaging gives every forecast equal weight; macro averaging gives every agent equal weight after computing its own score. Both can be valid, but they describe different populations. We should label the estimand explicitly when a few active native LLM agents generate most of the evaluation questions.
A two-agent ranking reversal
In an illustrative panel, agent X produces 90 questions and agent Y produces 10. Model A scores 0.10 on X and 0.50 on Y. Model B scores 0.20 on X and 0.20 on Y. A’s micro average is (90 × 0.10 + 10 × 0.50)/100 = 0.14, better than B’s 0.20. A’s macro average is (0.10 + 0.50)/2 = 0.30, worse than B’s 0.20. The arithmetic exposes a population choice rather than a contradiction.
Freeze agent eligibility
Define whether an agent needs a minimum resolved count and how failed or absent forecasts enter the report. Excluding agents with few observations can remove quieter mandates and new deployments. A macro score from one question per agent can be extremely noisy. Show count distributions, both summaries, and the exact weighting rule. Keep agent identities pseudonymous in public research artifacts where privacy requires it.
A repeated market question answered by many agents can inflate both agent and forecast counts without increasing independent market evidence. Preserve the shared question key and analyze time blocks or episodes. For model comparisons, pair questions inside each agent where possible. Report whether a model changes agent activity, because that can change the micro-weighted population. A fixed question panel isolates prediction behavior from endogenous invocation frequency.
From fixture to a held-out result
The scikit-learn scoring documentation provides the underlying methodological reference. The calculations and audit design here are illustrative extensions for an agent forecast record.
Continue with the evidence-dependence note for the related question of population weighting. Our benchmark card keeps this forecast-level comparison separate from action and execution results.
In a fleet comparison, publish the two weighted summaries beside the distribution of agent question counts. A large gap between them is information about heterogeneity. It should prompt inspection of lightly represented agents and shared market episodes. The choice of primary summary follows the research objective and should remain fixed before the model identities are revealed.