What We Do After a Directional-Edge Null

By DX Research Group · · DXRG findings

The historical fleet’s unfavorable return and signal results narrow the next experiment to measurable mechanisms and held-out prediction.

Neither fleet in our six-month record showed a directional edge. In the historical perpetuals fleet, cumulative realized P&L from June 8 to July 26, 2026 was -$217K at a common 5.5-basis-point fee rate and -$148K at zero fee. Removing fees left a substantial loss, so a fee-only explanation would send the next experiment in the wrong direction.

The continuous-record companion preserves these unfavorable findings alongside the pre-alpha system’s controls research. The larger fleet record runs through August 15 and includes mostly paper fills. P&L windows, model comparisons and signal probes each have their own populations; we keep those denominators separate.

Multiple checks point toward the same boundary

Against a matched sample of 1,961 Hyperliquid leaderboard traders, the fleet had a 41% roundtrip win rate against retail’s 50%, at nearly identical median holding times. Fifteen percent of week-active agents were net-positive against 53% of retail accounts. Leaderboard selection and the historical fleet’s fill assumptions limit how broadly this comparison travels.

Within the fleet, no lower day-clustered confidence bound exceeded zero across eight strategy postures, six prompt templates and nine cohorts. A roughly 77K-candidate signal ledger sorted at chance, while probes over 689 traces sat at or below their permutation null. A paired comparison with 770K tokens of context added no measured gain.

These observations reduce support for the tested directional hypotheses. They also stop us from attributing an attractive isolated subgroup to skill without checking its calendar footprint and uncertainty. The paper includes retractions for exactly that kind of inference failure.

Three distinct next decisions

We would first ask whether a candidate improves prediction on saved, held-out market questions. That requires a proper scoring rule, a declared target horizon and a baseline with the same information. A better narrative or longer context is evidence of changed model behavior; predictive improvement needs measured outcomes.

Second, we would test executable risk controls on identical entry paths. The record’s bracket replay offers a concrete example where a mechanical intervention changed simulated economics despite the directional null. Its implementation and live execution remain separate questions.

Third, we would compare inference cost when decision quality is tied. The 416-scenario replay league produced overlapping model-quality intervals, while reported inference spend was $8.25 for qwen3.7-plus and $203.61 for claude-fable-5. Those prices describe that league, and current procurement needs a fresh check. The result demonstrates why return skill and operational economics warrant different scorecards.

Progress we can defend

Our controls paper documents measured reliability interventions, including trace behavior rather than demonstrated return gains. That gives the research program useful work even after a market-skill hypothesis fails.

DXAP’s current homepage describes persistent agents, a tool harness and external policy checks. Those are platform criteria a reader can inspect, while a superiority claim about returns needs a matched current-product evaluation. We see frontier research as the ability to turn a clear null into a smaller, better specified experiment and preserve its unfavorable outcome if that experiment also fails.

Sources

Related field notes