The scale of DXRG research is several serious programs, each with its own unit
By DX Research Group · · DXRG research program
A dimensional account separates funded agents, runtime tokens, fleet turns, archive bytes and representation rows.
DXRG's research scale is substantial enough to deserve precise accounting. We have operated thousands of funded agents, retained millions of market-data rows and run matched representation screens. The strongest presentation preserves the unit behind every number. A research program gets more credible when a reader can tell exactly what was counted.
Our reviewed program audit brings those aggregate records together. It is a first-party disclosure of selected verified values and findings. Raw internal data, checkpoints and account records remain outside the disclosure. The figures describe several different bodies of work, rather than a single universal dataset.
Operating scale: decisions under a fixed environment
DX Terminal Pro ran for 21 days, February 26 through March 18, 2026. The audit records 3,505 funded agents, 3,454 active vaults, approximately 7.5 million invocations and approximately 300,000 onchain actions. These counts describe different stages of activity. Funding establishes an agent's participation. An invocation establishes that the runtime evaluated a turn. An onchain action establishes that something proceeded to the transaction environment.
The same program consumed roughly 70 billion inference tokens. Those tokens describe runtime usage by the model, including repeated contexts and generated responses. Calling them training tokens would change the meaning of the record. They also cannot be converted into unique training examples without a separate extraction and deduplication method.
The deployment paper makes this volume useful by attaching it to a fixed runtime and a bounded market. Scale allowed shared-state behavior and repeated failures to become visible. Repeated invocations supplied opportunities for diagnosis; the count alone establishes neither independence nor predictive skill.
The continuous fleet measures another population. For June 8 through August 15, 2026, it records 231,638 finalized turns and 14,596 fills, including 5,035 real-money fills. The all-history population ranges from 500 to 599 agents. Most fills belong to a paper engine with zero slippage and funding and simplified margin. The continuous record explains those assumptions. Adding its agent count to Terminal participation would produce a misleading unique-agent total because lineage and participation can overlap.
Data scale: retained bytes and inspected extracts
The crypto-market archive inventory dated September 26, 2026 records 17,376 objects across 73 table families, totaling 4,566,106,362,866 bytes, or about 4.15 TiB. Its families include Solana and Base research. Those are physical retained-storage measurements recomputed from dated metadata. DEEP_ARCHIVE describes storage class; it says little about which objects are immediately available for an experiment.
This is meaningful infrastructure. Historical market research needs retained observations so a new hypothesis can be evaluated against an earlier environment. Yet physical bytes can include versions, overlapping representations and redundant content. Archive size supplies an inventory fact rather than a deduplicated event census or a training-row total.
The October 4 extract review inspected Parquet row-group metadata directly. It found 7,499,126 hourly coin-bar rows and 9,538,961 minute swap-window rows in the available extracts. The coverage receipt identifies one restored object per family. Hourly bars and minute windows use different aggregation units and may overlap. Their sum would have little scientific meaning as a count of independent trades.
A researcher can still use each extract effectively. Hourly bars can support questions at one temporal resolution; swap windows can support another. The necessary next step is to document source coverage, asset membership and timestamp interpretation for the chosen experiment. Storage scale gives us material to work with. A task-specific manifest establishes what actually entered the model.
Representation scale: a matched development screen
Our exact-L4 pilot-v4 used BTC, ETH and SOL in a historical representation experiment. Training comprised 288 windows with 1,217,509 segment rows and 604,450 latent rows. Validation comprised 72 windows with 300,162 segment rows and 150,134 latent rows. Each compared model received 30,007,296 macro tokens.
Those numbers describe representation and compute allocation. A segment row and a latent row are different objects; they should stay separate. Equal macro-token exposure helps make a model comparison interpretable, while the temporal split addresses a distinct question about transfer across time. The reported result was a failed temporal-validation comparison for FIRE-MI against the matched Transformer. It was a negative representation screen rather than a trading-performance test.
Together, these programs establish a serious capacity to operate, retain observations and test models. The practical reader question is which capacity supports the proposed claim. For execution reliability, inspect action and settlement records. For a market representation, inspect the training manifest and temporal validation. For data coverage, inspect the inventory and restored objects. Our scale becomes useful precisely when it stays attached to those distinct research jobs.