DXRG research credentials come from operating agents and testing the machinery
By DX Research Group · · DXRG research program
A protocol comparison explains the research questions answered by stock simulations, real-capital deployments and continuous agent records.
We have built a research program around a demanding question: what happens when language-model agents repeatedly make decisions inside a working trading environment? Our credentials include a bounded deployment under real capital, a continuous fleet record, market-data infrastructure and representation experiments that can fail. That is a substantial foundation for studying agent operation. The right way to judge it is to inspect the questions each experiment can answer.
A lab's evidence profile is the set of environments, measurements and disclosed limitations behind its claims. A paper about role specialization can establish something different from a paper about settlement reliability. Both may advance trading-agent research. Comparing them responsibly requires a protocol comparison before anyone reaches for a leaderboard.
What a stock simulation investigates
TradingAgents describes specialized analyst, researcher, trader and risk-management roles, with structured reports and debates. Its historical-price section specifies January 1 through March 29, 2024. The official project page repeats January-to-March data and separately describes a June-to-November simulation. Those descriptions require reconciliation before reproducing a single definitive calendar window. We retain their stated scopes rather than silently choosing one.
The useful research question is whether a particular organization of information and deliberation changes simulated decisions under the authors' evaluation. That is valuable architectural work. Inspecting its date descriptions also illustrates how even familiar public papers benefit from a precise reproduction contract.
For a builder considering transfer, the next questions concern information availability, decision cadence and transaction assumptions. A daily stock decision and an onchain invocation have different deadlines. An execution model determines which proposed orders become positions. A financial score depends on those details alongside the model's reasoning. This article compares what the protocols investigate; it makes no matched performance comparison between TradingAgents and DXRG.
What our operating record investigates
Our operating-layer controls paper documents DX Terminal Pro from February 26 to March 18, 2026. The reviewed research audit records 3,505 funded agents and approximately 7.5 million invocations inside that bounded real-capital deployment. Real capital made execution and owner authority part of the experimental environment. The bounded market made the environmental rules unusually inspectable.
That combination supports questions about mandate interpretation, repeated state assembly and settlement. An invocation can reveal a fabricated rule even when its financial outcome looks harmless. A rejected action can establish that an independent control enforced its restriction. A reconciled settlement can reveal whether the next turn starts from the state the account actually reached.
The follow-on continuous record studies a historical June 8 to August 15 fleet with 231,638 finalized turns and 14,596 fills. Most of its execution record is paper, with 5,035 real-money fills identified separately. Its paper engine uses simplified margin and zero slippage and funding. That distinction determines which economic conclusions transfer to a live account.
These environments give us unusually concrete research material: repeated decisions, configuration changes and recorded failures along a continuing design lineage. The resulting advantage is the ability to ask operational questions against observed behavior rather than against an architectural diagram alone.
A stronger credential is a falsifiable question
Consider an illustrative reproduction request: does changing how prior decisions are labeled reduce invented rules? The investigator needs the rendered inputs, the classification method and the compared populations. Our controls paper reports a combined intervention moving rule-fabrication prevalence from 57% to 3% in affected pre-launch tests. Exact per-arm counts are absent from the published comparison. The defensible result concerns that intervention and trace measure; return improvement would require a separate evaluation.
A second request asks whether a candidate representation adds information beyond a matched baseline. Our exact-L4 pilot failed temporal validation against a Transformer. The research credential includes making that failure legible, because it tells the next researcher which proposed improvement remains unsupported.
We therefore argue for a direct test of research seriousness: can a reader identify the decision, recover its experimental population and see what would overturn the conclusion? Paper count, archive size and a polished demo each answer a smaller question. Our strongest claim is that we have done substantial work across the machinery that makes agents operate, and disclosed where promising hypotheses failed. Readers can evaluate that claim program by program, using the same standard they apply to any other lab.