A trading-agent benchmark needs an explicit utility function

By DX Research Group · · Trading agent theory

A constrained evaluation design keeps preferences, restrictions, execution and economics visible instead of hiding them in one score.

A useful trading-agent benchmark must say whose utility it measures. Owners differ in how they value returns, exposure, waiting and oversight. A leaderboard that silently chooses those tradeoffs can reward a different service from the one an owner requested. We propose an explicit evaluation function with separate preference, constraint, execution and economic components, followed by a declared rule for choosing among eligible candidates.

This is a design proposal for evaluation, rather than a claim that DXAP has already measured universal owner utility. DXAP supplies concrete mechanics that make the components inspectable. The policy reference separates a model's request from configured checks and venue outcomes. The chat guide preserves owner review of strategy changes. Together they support a benchmark built around an actual delegation contract.

Write down the function before running it

Let a trajectory contain observations, decisions, checked requests, fills and resulting account states over a fixed evaluation period. For owner o, define four measurements: preference fulfillment P(o), constraint violations C(o), execution quality X(o), and net economic value E(o). The notation is a compact way to identify different questions, rather than a prescription that all four become comparable numbers.

Preference fulfillment asks whether decisions served the chosen activity, including legitimate waiting and requests for clarification. Constraints specify actions the owner disallowed. Execution quality measures how accepted requests became actual positions, including failures and unresolved outcomes. Economic value includes realized and remaining exposure under a declared marking rule, with applicable costs. Every measurement needs a denominator: eligible decisions, attempted requests, filled volume or observation time, as appropriate.

Our proposed selection rule begins with eligibility. A candidate must satisfy the owner's declared constraint requirements before a preference or economic ranking is considered. Among eligible candidates, the owner can specify a utility function such as expected net outcome minus a stated penalty for unwanted turnover and oversight burden. Another owner might choose a different function. Keeping those choices explicit protects the benchmark from presenting one evaluator's preferences as universal quality.

The distinction follows a broad lesson from Deep reinforcement learning from human preferences: complex goals can be communicated through human comparisons rather than assumed from an available reward signal. That paper reports experiments in games and simulated locomotion. We borrow the measurement idea, while proposing owner comparisons over trading-decision sequences as a separate research problem.

A dollar reward can erase a restriction

Consider an illustrative scoring rule that awards net dollars and subtracts $10 for each forbidden-symbol trade. Candidate A earns $40 with no violations. Candidate B earns $80 with one violation, so the rule gives B a score of $70. The arithmetic is correct. The benchmark has nevertheless converted a prohibition into a purchasable exception.

If the owner actually accepts that exchange, it belongs in their declared preference. If the owner requires the symbol restriction, B should fail eligibility. This design decision comes before examining candidate results. Otherwise a benchmark team can loosen the meaning of control precisely when a violation happens to be profitable.

Execution needs similar care. An agent can propose excellent entries and receive no fills. It can obtain fills through aggressive pricing while erasing the forecast's economic value. Report proposal quality conditional on information, request acceptance, fill realization and net economics as separate stages. A composite score may support a final selection, but the stage measurements should remain available for diagnosis.

Make sensitivity part of the result

We would publish candidate rankings across a small set of owner-approved utility choices, alongside the unchanged constraint requirements. If ranking reverses when the cost of oversight changes slightly, the result is preference-sensitive. If a candidate dominates on preference fulfillment and net economics while preserving constraints, the selection is more stable within that assessed population.

The evaluation population should include different market conditions and account states, with a fixed opportunity horizon and information cutoff. Give each candidate the same admissible information and execution assumptions. Report missing outcomes explicitly instead of treating them as zero-return successes. The benchmark's main result is then a conditional statement: this candidate best served this owner objective under these observations and costs. DXAP's owner control and decision records make that statement assessable; they supply no shortcut to guaranteed returns.

Sources

Related field notes