Our failed hypotheses make the DXRG research program stronger

By DX Research Group · · DXRG research program

Specific retractions and development audits show how rejected claims improve the next experiment.

A research lab earns credibility by making its strongest attractive claim survive scrutiny, and by removing it when it fails. We have done both in the DXRG program. Our published record includes retractions, an explicit directional-edge null and a failed representation screen. Those findings improve the next experiment because they identify the error a future claim must overcome.

The useful question for a reader is simple: does a negative result change what the lab does? A disappointing score stored in a forgotten notebook changes little. A documented failure that alters timestamps, validation or execution design becomes a durable research asset.

Three retractions with different failure mechanisms

The continuous agent record describes three earlier results that were retracted. One apparent HIP-3 lag edge proved to be a timestamp artifact. Two apparently net-positive findings failed calendar-footprint checks. The distinction matters because the repair differs.

A timestamp artifact changes which information was available when a decision could be made. A historical relationship can look predictive because an observation has been assigned to the wrong moment. The corrective test reconstructs event time and decision-time availability. A better model fitted to the same mistaken chronology would deepen the error.

A calendar-footprint problem asks whether the strategy's apparent advantage comes from the days on which it participated. If a selected cohort traded during a favorable interval, its aggregate outcome can reward that calendar exposure rather than its decision rule. Calendar-matched comparisons and day-level uncertainty address the mechanism. Pooling more trades from the same favorable period can make the misleading result look more precise.

These retractions give us a concrete standard for new work: chronology and participation must be visible before a strong economic claim is accepted. The continuous record's methodology canon follows from such failures. Its principles include using the market day as an inferential unit, reporting a full sweep when a band was searched and requiring an increment over a concrete baseline. The discipline improves the question as much as the statistical calculation.

A representation experiment that failed its intended test

The reviewed research-program audit records an exact-L4 pilot-v4 development screen on BTC, ETH and SOL. FIRE-MI failed temporal validation against a matched Transformer. Both compared models received 30,007,296 macro tokens. The outcome concerns representation research under that pilot, with the temporal comparison preserved.

The benefit is a sharper allocation decision. A novel representation can deserve investigation without deserving adoption. Once its tested increment fails, the next experiment needs a specific changed mechanism or a genuinely different evaluation question. Repeating the same flattering development summary would spend compute while leaving the rejected premise intact.

There is also a general lesson about baselines. A sophisticated architecture can improve a chosen metric while contributing less than a simpler matched alternative. Giving both models comparable exposure makes the failure informative. Our claim to research seriousness includes reporting that comparison plainly.

ECTO shows why accuracy needs an audit

The same disclosure reviews ECTO's historical Solana state-lattice transformer development. On 12,982 development examples, recorded accuracy was 87.26%, while the flat-majority baseline was 86.35%. The lift was about 0.91 percentage points. That calculation immediately changes what the headline means: most examples belonged to the flat class, so a high accuracy percentage alone gives little evidence of useful event discrimination.

The audit also identifies checkpoint selection using pump precision on the test partition, uncertain consistency of mint membership between splitting methods and overlapping windows. Those are methodological findings with direct consequences. A selected checkpoint needs a separate untouched assessment. Exact saved memberships are needed to establish asset separation. Dependent windows require uncertainty tied to event clusters rather than a fiction of independent samples.

ECTO's label adds another issue: its reference price begins at input-window start. Some target movement is therefore observed before decision time. A target-label score must be realigned before it can support a decision-time forecast. A later 13-example review across five tokens was entirely flat in truth and prediction; its 100% accuracy contributes no pump or dump discrimination evidence.

Our next contract follows from those observations: persist split membership, separate development and calibration from untouched chronological evaluation, compare class-specific baselines and replay executable actions with costs. These are proposed requirements for the next study, rather than results already achieved.

A lab that retracts a result loses the right to repeat that result and gains a better description of the research problem. We see that trade as essential. Our strongest research asset is a growing record that says which mechanisms held, which measurements misled us and exactly what the next claim must demonstrate.

Sources

Related field notes