ECTO accuracy changes meaning when the flat baseline is visible

By DX Research Group · · DXRG research program

An 87.26% development accuracy becomes a sharper research result when class prevalence and threshold support are reported.

Our ECTO development evaluation learned substantial class structure, but its headline accuracy sits close to a very simple comparator. Across 12,982 recorded examples, accuracy was 87.2593%. Predicting flat for every example would have achieved 86.3503%. The difference is 0.9090 percentage points. That comparison changes the research question from whether a high percentage looks impressive to which predictions improve on the dominant class.

We disclose those values in the DXRG research-program audit. The evaluation belongs to historical Solana token development, with five-minute inputs and fifteen-minute target windows. Model selection used the evaluated partition, and overlapping windows make examples dependent. These properties define how much weight the recorded metrics can carry.

Start with the population

The actual class counts were 11,210 flat examples, 982 pump examples and 790 dump examples. The flat count divided by 12,982 gives the 86.3503% majority baseline. Together, the two movement classes account for 1,772 examples, or about 13.65% of the recorded population.

The reported accuracy corresponds to 11,328 correct predictions. The all-flat comparator gets 11,210 correct. ECTO therefore records 118 more correct classifications overall. This arithmetic makes the aggregate improvement tangible without implying that the two methods make the same errors.

They plainly have different behavior. The majority comparator finds none of the pump or dump cases. ECTO's pump recall was 80.35%, and dump recall was 84.68%. Those recalls show that the fitted classifier identified many movement labels in the development examples. The accompanying false positives decide whether that discrimination is useful.

Pump precision was 62.08%, while dump precision was 40.16%. In the dump class, fewer than half of the 1,666 predictions matched the dump label. High recall and lower precision can coexist: a classifier catches many actual cases by making a broad set of predictions. Accuracy averages that tradeoff together with the much larger flat class.

The class-specific results are therefore more informative than a single verdict about the model. We learned that ECTO's development behavior was neither equivalent to constant-flat prediction nor adequately described by the high aggregate accuracy alone. It traded some flat correctness for detection of movement classes.

A confidence threshold trades support for precision

The recorded pump threshold results make another tradeoff visible. At a threshold of 0.3, precision was 47.23% over 1,857 pump predictions. At 0.5, precision was 76.56% over 836 predictions. At 0.7, precision reached 88.17% over 169 predictions.

For the last threshold, 88.1657% corresponds to 149 matching pump labels among 169 predictions. Compared with the 0.5 threshold, the higher threshold retains about one fifth of the predictions. The apparent improvement in precision comes with much narrower coverage. A reader choosing between these settings needs both columns.

Those 169 predictions also cannot be treated as 169 independent market opportunities. The input windows move at one-minute strides despite covering five minutes, and the fifteen-minute targets overlap too. Nearby high-confidence predictions may describe the same price episode repeatedly. Counting rows measures evaluation support; counting independent episodes requires grouping the records.

A conventional uncertainty calculation that assumes independent rows would overstate how much evidence those predictions contain. We would evaluate uncertainty by event clusters and show the number of distinct episodes and assets alongside the threshold result. That is a proposed repair to the measurement, with no new interval claimed here.

The comparator must fit the question

For aggregate classification, the all-flat baseline is essential. For pump detection, precision and recall against pump prevalence give a more useful starting point. For probability quality, calibration and proper scoring need their own assessment. For a trading decision, executable prices and costs enter a further experiment.

Our next evaluation should freeze threshold selection in development, then compare the selected model on untouched chronological and asset partitions. It should retain the full class counts and report movement-class behavior even when overall accuracy falls. A later recorded validation contained only 13 examples across five tokens, all actually and predictively flat. Its 100% accuracy adds no pump or dump discrimination evidence.

The constructive result of the audit is a clearer target for research. ECTO displayed movement-class discrimination during development, while its total correctness improved only modestly over a constant prediction. We can now ask whether the class-specific gains survive independent evaluation, and whether the high-confidence subset contains enough distinct events to support the next experiment.

Sources

Related field notes