Choose a model with a latency and validity budget
By DX Research Group · · DXAP platform
A model review based on timely valid decisions rather than reputation alone.
Choose the model that can complete your required decision task within a useful information window, with an acceptable rate of valid outputs. Model reputation alone cannot settle that operational choice. We would begin with the strategy's deadline and evidence requirements, then compare behavior under a shared harness.
The configuration reference describes selecting a model and evaluation interval. The historical continuous record reports replay comparisons whose economic differences remained within uncertainty, alongside substantial inference-cost differences. Those historical results are evidence about that replay and its models, rather than a ranking of today's available product choices.
Define the deadline before the contest
Illustrative comparison: an owner has eight saved review cases, each with identical strategy, account state, and tool evidence. They decide that an answer arriving after 90 seconds misses the particular event-review deadline. Model A returns seven valid answers, six within the deadline. Model B returns eight valid answers, five within the deadline.
Under this deliberately narrow acceptance definition, A produces six timely valid answers out of eight, or 75%. B produces five out of eight, or 62.5%. Those small samples support a case-level comparison; a stable population estimate requires broader evidence. Record all failures and late responses rather than averaging latency only across successful cases.
The 90-second cutoff is an owner-chosen test requirement for this fictional use case. A slower position-management strategy could use another deadline and reach another conclusion. Make that source of the threshold explicit.
Define validity through the action
A valid response should use the allowed instrument, correct account state, required evidence, and applicable strategy version. If it proposes an action, that action should be interpretable and admissible under the configured policy. A fluent paragraph that invents evidence fails this review even if it arrives quickly.
Waiting can be a valid answer. Score whether waiting follows the mandate, rather than treating every abstention as a model failure. Conversely, an always-wait policy can look operationally clean while offering little decision value. Our evaluation guide provides the broader dimensions needed before claiming useful trading performance.
Keep platform price and research cost separate
DXAP's published alpha pricing includes no separate model charge, according to the current fees reference. Internal inference economics from an old experiment belong in the historical research accounting, separately from the current owner's published invoice terms. For the owner, the relevant cost can include delayed decisions, missing evidence, and behavior that changes turnover.
Review a model switch as a proposed configuration change. Compare later records with the settings actually active at the time and attribute any simultaneous strategy revision separately. A good initial choice is the one whose observed behavior fits the required task. Revisit it when the strategy, model version, or tool environment materially changes.