A timestamp-artifact retraction changes the research contract

By DX Research Group · · DXRG findings

Use the withdrawn HIP-3 lag claim to specify what a point-in-time audit must preserve.

The continuous record paper retracts an earlier HIP-3 lag-edge claim because it turned out to be a timestamp artifact. We should retain that retraction as a method finding. It shows that a result can appear predictive when its information timing is wrong, even before any debate about model quality begins.

The continuous record companion, updated September 9, 2026, says three earlier claims appear only as retractions: the timestamp-related lag claim and two net-positive claims invalidated by calendar-footprint checks. The complete raw reconstruction of the timestamp failure remains unavailable in the public companion. We therefore keep the precise defect at the reported level instead of inventing which feed or timestamp field caused it.

Reconstruct the information boundary

An illustrative quote occurs at 12:00:00, reaches the data process at 12:00:04 and enters an agent prompt at 12:00:06. That quote was unavailable to a 12:00:02 forecast under the receipt-time contract. Joining by event time alone would place it before the decision and create four seconds of unavailable information. This fixture illustrates a generic failure mechanism; the actual HIP-3 implementation needs its own reconstruction.

A concrete audit artifact would preserve event time, receipt time, persistence time and prompt-inclusion time for every relevant input. It would state which clock determines availability and record timezone and timestamp precision. A future-data assertion should then reject any input that arrives after the decision boundary under that contract.

The audit also needs a frozen prediction target and outcome window. Repairing the timestamp join changes the eligible examples and may change apparent lead-lag relationships. A result should be rerun with the corrected availability rule and compared with a reference that has the same information boundary. Until that result exists, the retracted edge stays withdrawn.

A retraction is a scope correction

The paper companion reports a rolling p-value moving from 0.0067 to 0.19 to 0.0277 before a claim was called. Those values illustrate why a selected snapshot can flatter a fragile result. Quantifying the defect and its prevalence across trading studies requires additional measurement. Timing validity and selection history require their own artifacts.

Our controls paper companion describes invocation-level traces linking rendered prompts to actions in an earlier 21-day, twelve-token Base deployment. That trace structure suggests a practical place to preserve information timestamps. Correctness of later feed joins requires feed-specific checks.

For a reader evaluating an agent research claim, a useful request is the availability contract and a boundary fixture that demonstrably fails when future information is inserted. Public paper figures remain useful for understanding the program's correction. A live trading advantage needs a new, time-valid result with its own holdout, costs and uncertainty.

Sources

Related field notes