Research

We publish deployment records, engineering methods, evaluation protocols, and reusable data for people building and assessing AI trading agents.

Research programs and deployments

These projects connect first-party deployments to the harness, runtime, controls, and evals we use around the model.

DX Terminal

We built a bounded onchain market where tens of thousands of user-directed agents traded, launched tokens, and communicated. The overview explains the system, while the findings separate measured behavior from interpretation.

DX Terminal Pro

We ran a 21-day real-capital deployment on Base and preserved the path from user instruction through execution and settlement. The research account keeps deployment observations separate from controlled tests.

Trading-agent harness and runtime

Our architecture begins with an authenticated mandate and carries typed actions through policy validation, execution, settlement, reconciliation, and trace-based evaluation.

Trading-agent evaluation registry

We publish benchmark cards, harness-transfer tests, state and memory fixtures, and versioned data so readers can inspect the method and evidence class behind each result.

Public datasets

Each versioned entry lists its formats and the articles that present it. Null and unrun fields remain explicit until a registered evaluation produces evidence.

Dataset
Formats

DXRG Trading-Agent Prompt-Compilation Ablation Registry

Eight unrun comparison templates for testing prompt-compilation changes while freezing the model, state, memory, action schema, policy, execution adapter, and evidence class.

Version 1.0.0Published 2026-07-31

DXRG Trading-Agent Harness-Transfer Evaluation Card

Eight controlled comparison templates for separating model, prompt compilation, state, memory, action schema, policy, execution, and environment effects in trading-agent evaluation.

Version 1.0.0Published 2026-07-28

DXRG Trading-Agent Trace-Derived Regression Registry

Eight versioned cases for turning linked trading-agent traces into bounded regression tests with frozen components, named interventions, diagnostics, downstream checks, and evidence-class limits.

Version 1.0.0Published 2026-07-24

DXRG Trading-Agent Execution and Reconciliation Test Matrix

Eight failure fixtures for final-payload binding, request identity, ambiguous timeouts, acknowledgements, partial fills, fees, reconciliation, and controlled recovery.

Version 1.0.0Published 2026-07-24

DXRG Trading-Agent State and Memory Test Checklist

Eight failure fixtures for evaluating state identity, freshness, portfolio reconciliation, order lifecycle, venue state, memory provenance, action binding, and post-settlement feedback.

Version 1.0.0Published 2026-07-22

DXRG Trading-Agent Mandate Compilation Evidence Table

Each record states the evaluation design, intervention or setting, fixed components, observed measure, interpretation, and evidence boundary. Controlled pre-launch interventions and historical live setting gradients remain separate evidence classes.

Version 1.0.0Published 2026-07-21