When should a failed trading-agent theory be retired?
By DX Research Group · · Frontier research
A proposed research decision record distinguishes a tested failure from an unresolved or untested mechanism.
A failed theory should be retired when its defining prediction has received a fair, discriminating test and the remaining uncertainty no longer justifies the next experiment. We propose recording that decision explicitly. This is an unrun research protocol, with stopping criteria chosen for a particular question rather than imposed as a universal rule.
The continuous record companion reports a historical directional-edge null alongside concrete harness failures. That combination matters: a failed directional claim can coexist with useful engineering work. The controls paper companion reports bounded trace improvements whose economic implications remain separate. A theory about better forecasting needs forecasting evidence; a theory about fewer invalid actions needs action evidence.
Write the prediction that could fail
Suppose an illustrative theory claims that a new event summary improves a fixed-horizon probability forecast because it preserves surprise information lost by the existing summary. Its discriminating test compares both summaries from the same event packet, with identical arrival time, model and forecast question. A placebo summary tests whether extra text alone explains the change. A direct extraction check determines whether the supposedly preserved fact actually reaches the model.
If extraction fails, the mechanism was inadequately delivered. If extraction succeeds but forecast scores do not improve, the claimed predictive mechanism has received a more informative challenge. If forecasts improve but a fixed trading policy loses after costs, the forecasting theory and the economic theory receive different statuses. Keep those distinctions in the decision record rather than collapse every failure into a single verdict.
For illustrative budgeting, a follow-up experiment costing $2,000 has a 10% assessed chance of resolving a decision worth $5,000 to the research program. Its simplified expected decision value is $500 before cost, yielding minus $1,500 after cost. This calculation depends on subjective inputs and omits other learning value. It clarifies which assumptions would justify continuing, instead of treating an attractive hypothesis as a reason for unlimited retries.
Four outcomes for the research record
Record a tested prediction as supported within scope, rejected within scope, unresolved or untested. A scope-specific rejection can retire the immediate development claim while leaving broader variants open. An unresolved result needs the source of uncertainty: sparse events, noisy labels or an intervention too weak to distinguish mechanisms. Repeating the same weak test adds less information than repairing that source.
We would attach the theory's original statement, its strongest unfavorable artifact and the next experiment's expected information to one dated decision. Include the cost already spent and avoid allowing sunk cost to stand in for future value. If a new mechanism is proposed, register it as a new theory with a new discriminating prediction.
Retirement then becomes a useful research output: a bounded claim with its evidence and remaining questions preserved. It frees effort for repairs or alternative hypotheses while preventing a future team from quietly reviving the same failed claim as if its earlier test never happened.