ECTO taught us that a split must preserve membership

By DX Research Group · · DXRG research program

A development audit explains why a shared random seed and a held-out label leave two separate independence questions unresolved.

We learned two specific lessons from auditing ECTO's evaluation history: save the actual partition members, and separate checkpoint selection from final assessment. These are practical improvements to the research system. They determine whether a promising classifier has faced a genuinely new population or merely another view of the examples used to develop it.

Our research-program audit records both issues. Training shuffled a list derived from a set of sequence mints, while later evaluation shuffled a chart-file list. Training also selected the best checkpoint using pump precision on the test partition. Each issue affects a different part of the claimed held-out result.

The correct scope is historical ECTO development on Solana token microstructure. Its recorded evaluation includes 12,982 examples with five-minute input windows and fifteen-minute targets. The audit allows us to preserve the measured behavior while narrowing the interpretation and designing a cleaner successor experiment.

A seed reproduces an operation on a particular list

A random seed is useful because it can reproduce a shuffle given the same algorithm and the same ordered input. It cannot establish that two independently assembled lists contain identical members in identical order. The starting population is part of the split definition.

Here is an illustrative example. Suppose the training script starts with token identifiers in the order A, B, C, D. An evaluation script starts with C, A, D, B. Even if the seeded shuffle applies the same positional permutation to both lists, the token occupying each resulting position can differ. Selecting the first half after the shuffle can therefore assign different token members to the evaluation partition.

The ECTO audit found an additional reason to preserve that distinction: one list came from sequence mints and the other from chart files. Those are different routes for constructing membership. A common seed supplies no proof that the resulting partitions are identical or disjoint from training.

This finding establishes an unresolved membership question, rather than proving that a particular token leaked. Certifying disjointness requires the actual training and evaluation identifiers, followed by an explicit intersection check. A trustworthy receipt would also account for missing members and duplicates so a changed file inventory cannot quietly change the population.

We would save those memberships when the dataset is created. Later evaluation would load the saved assignment instead of reconstructing it from whatever files happen to be present. That turns a claim about intended partitioning into an inspectable statement about the examples actually used.

Model selection changes the role of a partition

The second issue remains even if membership is repaired perfectly. ECTO training chose the best checkpoint using pump precision on the test partition. The evaluated labels therefore influenced which model was retained.

An illustrative run makes the consequence clear. Imagine ten epochs produce ten checkpoints, each scored against the same evaluation labels. Selecting the checkpoint with the highest pump precision rewards both real learning and favorable variation on those labels. The selected score summarizes a model-selection procedure that has seen the evaluation result ten times. Renaming that partition cannot reverse those decisions.

There is nothing inherently wrong with using a partition to choose an epoch. That is a development function. The final assessment needs a further population whose outcomes played no role in choosing the checkpoint, threshold or representation. Our audit names the original result according to its actual use: development and model selection.

The two repairs complement each other. Saved membership answers which assets and examples entered each partition. An untouched final assessment answers whether the retained model generalizes after the researchers finish choosing it. Either repair alone leaves the other question open.

A successor experiment with an inspectable boundary

Our next research contract separates development, calibration and untouched chronological and asset evaluation. The development population supports architecture and epoch choices. Calibration supports probability or threshold adjustments under its stated protocol. Final assessment begins after those decisions are frozen.

The windows require temporal care as well. Five-minute inputs and fifteen-minute targets overlap across one-minute strides. A timestamp boundary must account for the interval used by each example, and uncertainty should reflect event dependence. Saving token membership alone cannot resolve overlap among windows from the same episode.

We would publish aggregate receipts for the resulting partitions, selection rule and evaluation population, retaining private raw records internally. Readers could then see whether the comparator, label and threshold stayed fixed before final evaluation. The historical ECTO numbers remain useful for understanding what the model learned during development. The successor experiment's contribution would be the stronger claim that survives after both membership and model-selection independence are established.

Sources

Related field notes