Inside ECTO: a state lattice for token microstructure
By DX Research Group · · DXRG research program
What a 65,536-cell market representation and a 256-cell selection teach us about building a microstructure transformer.
We built ECTO to ask whether a structured representation of token activity could make short-window market behavior easier for a transformer to learn. The historical Solana research configuration organizes activity across price, volume, time and direction. Its distinctive engineering choice is the representation before the model: a state lattice with 65,536 possible cells and a selection of 256 cells for the model input.
That is a concrete architecture readers can inspect. Our aggregate research audit discloses the configuration and the development findings. The useful question is which market information survives this representation, and which distinctions disappear before attention ever sees them.
The arithmetic of the representation
The configuration uses 32 price bins, 32 volume bins, 32 time bins and two direction bins. Multiplying these dimensions gives 32 × 32 × 32 × 2 = 65,536 cells. A cell locates activity within that joint space. Price and volume alone would provide 1,024 combinations; the time and direction dimensions expand the representation by a factor of 64.
The disclosed price-range parameter is 0.1, and the log-volume range runs from -6 to 6. The input window covers 300 seconds. Those choices fix the coordinates in which activity is represented. They make the research reproducible at the configuration level while leaving separate implementation questions, such as exact boundary assignment, to a fuller representation specification.
ECTO selects 256 cells from the lattice. That is 0.390625% of the possible cells, or one cell per 256 possible locations. This calculation describes the selection relative to the full space. It says little about the fraction of observed market activity retained, because many possible cells may contain no activity and selected cells may contain very different amounts.
The transformer configuration has a model width of 256, four layers and eight attention heads. The number 256 appears both in cell selection and model width, with different meanings. One limits the selected representation; the other specifies the width of the model's internal vectors. Conflating them would obscure where information is discarded and where computation happens.
What the compression makes possible
Consider an illustrative five-minute token window containing repeated small trades and a few large trades. A raw event sequence and a lattice present different problems to the model. The raw sequence preserves event order directly. The lattice groups activity by its position in the configured dimensions, making repeated occupancy of a region easier to summarize.
That can be a useful inductive choice. The model receives an organized market state rather than being asked to discover every grouping from an arbitrary event list. Time bins preserve a coarse temporal coordinate, while the direction dimension separates two configured directions. The architecture gives us explicit representation choices to test.
The cost is equally explicit. Two windows can map to similar selected cells while differing in event order inside those cells. Activity outside the selected set can disappear from the model's view. Sparse unusual behavior may matter even when it contributes little to the selection criterion. These are questions about the encoder, upstream of whether a larger transformer would fit the retained data better.
We learned to treat representation and prediction as separate experiments. A fitted model can only use distinctions that survive the input transformation. Strong attention over a compressed input cannot recover an event that the input removed.
The next comparison belongs before a bigger model
A useful follow-up would hold the population, labels and split fixed while varying the representation. Compare the 256-cell selection with a larger selection and with a sequence baseline under a stated compute budget. Measure class-specific prediction changes alongside retained activity and representation collisions. This is our proposed experimental design, rather than a completed superiority result.
The time contract also deserves its own line in that comparison. ECTO's recorded input covers five minutes and its target window covers fifteen minutes. Those are the historical research horizons, not hours-to-days forecasts. The label references the input's opening price, which requires a separate decision-time audit before execution interpretation.
Our development work establishes that this lattice was a real architectural choice, with disclosed dimensions and model configuration. It also gives us a focused research agenda: determine whether grouping activity improves generalization after accounting for selection loss and evaluation validity. The contribution is a testable representation with identifiable failure modes. Its value will be decided by matched experiments that show which market distinctions it preserves well enough to predict unseen outcomes.