Randomizing candidate lists to measure trading-agent selection bias
By DX Research Group · · Frontier research
A proposed randomized render experiment tests how visibility and position affect asset selection independently of market information.
Which assets an agent sees can determine which assets it trades. Our continuous-record paper companion reports that 46.5% of historical entries used rendered mover symbols against an 8.9% random-availability baseline. Its render-boundary analysis is unusually informative because nearby symbols had similar market state yet different visibility. We would extend that observation through a PROPOSED randomized experiment, rather than assume a larger candidate list improves decisions.
Separate inclusion from placement
The experimental unit would be a saved agent turn. Each turn receives the same portfolio, mandate, timestamp and eligible asset universe. We first randomize which eligible symbols appear in a fixed-length shortlist, using a recorded inclusion probability for each symbol. We then randomize their order independently. This creates two identifiable questions: does display change selection, and does a higher display position change selection among displayed symbols?
A market-quality score used for analysis must be calculated from information available before the turn. We would stratify randomization by that score and liquidity to avoid accidentally replacing a liquid shortlist with an unusable one. The full eligibility rule remains identical across arms. An asset excluded by the user's mandate stays excluded regardless of its assigned experimental position.
Build one fixture that exposes the mechanism
An illustrative replay contains twelve eligible assets, of which six are rendered. A particular asset appears first in one assignment and sixth in another, with all its market fields unchanged. We save the model response, selected symbol, no-trade decision and typed action. A switch toward the first position reveals presentation sensitivity on that fixture. Its eventual return answers a different question and must be scored under the same forward window and execution assumptions.
Across the corpus, we would estimate selection probability conditional on visibility and position, accounting for each symbol's randomized inclusion probability. Day-level uncertainty matters because multiple agents may react to the same market shock. We would publish selection shifts even if every performance comparison remains inconclusive. A statistically clear presentation effect can coexist with an economically negligible change.
What would justify a new shortlist?
The decision criterion should include diversity, mandate compliance and net opportunity quality. Uniform random ordering might reduce position bias while making important information harder to find. A fixed ordering might improve interpretability but amplify a weak ranking signal. The experiment should reveal that tradeoff rather than assume one preferred layout is correct.
Our operating-layer controls paper shows that rendered instructions can materially change traces. Its fee-sentence intervention measured citation behavior, so we would keep trace changes distinct from return changes here as well. DXAP describes a data index feeding model proposals inside policy checks. That makes candidate rendering a concrete part of agentic trading engineering: a surface we can specify, perturb and evaluate as the harness evolves, without treating an attractive shortlist as proof of investment quality.