A directional null narrows the claim and improves the research agenda
By DX Research Group · · Market predictability
DXRG’s continuous record rules out easy claims about its historical fleets. Its strongest use is a precise domain of failure, not a universal impossibility theorem.
A well-described directional null gives researchers a concrete place to stop repeating an unsupported claim. It identifies the system, information and evaluation under which apparent market insight failed to become measured skill. We regard that precision as a research asset: it makes the next hypothesis explain what has changed and why the earlier result should cease to apply.
Our continuous record covers a 21-day real-capital Base tournament and a later Hyperliquid fleet observed from June 8 through August 15, 2026. It reports no demonstrated directional edge in either fleet. The later system used live prices, but most fills were paper with zero slippage and funding; those execution assumptions limit economic interpretation. This is evidence about historical agents and their operating conditions.
The finding has force because it constrains familiar explanations. The record includes failed signal and representation probes, rather than relying only on an unfavorable aggregate profit line. It also documents behavior shaped by configuration and rendered choices. An agent can produce consistent actions and detailed reasons while its directional information remains weak. Consistency and articulation therefore need their own assessment.
Define what a new result would have to escape
The useful object is the domain of the claim: a population of instruments, an information set, a forecast target, a runtime and an assessment window. A null in that domain carries directly into a proposal that changes only the marketing description. It carries less directly into a proposal that changes a substantive component, provided the new component has a plausible route to information or decision improvement.
Suppose a candidate adds a genuinely earlier source of event information. Its research case should show when that source became available, which decisions could use it and what incremental forecast target it serves. Suppose instead that a candidate adds a better stop implementation. Its case concerns the consequences of executing existing decisions. Both can be worth testing, but they challenge different parts of the historical result.
Changing the model name alone leaves a larger explanatory burden. A new model might extract useful information that an earlier model missed. The comparison must nevertheless show an increment on the same saved information before a broader live study attributes gains to model intelligence. Changing the forecast target at the same time would make the source of improvement ambiguous.
This gives us a practical claim map. A replication asks whether the historical null persists under comparable conditions. A forecast-extension study asks whether additional time-valid information helps. A control study asks whether the same uncertain forecasts produce safer or more efficient decisions under different execution mechanics. A transfer study asks whether a measured mechanism survives another population. These questions produce different receipts even when all belong to one agent program.
Failure to establish skill still leaves uncertainty
A confidence interval that includes zero can contain a small useful effect, a small harmful effect or both. Calling that result equality would require an additional equivalence question and a justified margin. Likewise, selecting the best subgroup after inspecting many candidates can create an impressive estimate with little fresh evidence. The null should survive into the report with its interval and selection history intact.
A universal impossibility claim would need to range over information and mechanisms the fleets never tested. Market-making, privileged information and latency advantages present distinct mechanisms; their existence or absence cannot be decided by a mid-horizon public-context LLM result. Even within public information, another target such as conditional risk can have a different answer from price direction.
The positive consequence for DXRG is a sharper allocation of research effort. We can demand that each proposed successor name the boundary it changes, while using the historical record as a replication reference. A public result that keeps unfavorable evidence accessible gives readers a way to judge progress rather than compare promises.
The next strong claim should therefore contain two linked facts: the earlier domain where skill failed to appear, and the newly measured mechanism that changes the result on untouched evidence. Until that second fact exists, the null remains the useful published finding. It preserves both the possibility of future improvement and the cost of pretending improvement has already happened.