DXRG research library
Agentic trading field notes
The frontier of agentic trading runs through the full decision system: what an agent reads, how it reasons, what it can execute, and what the record proves. These notes turn those questions into methods, examples, and testable theories.
Read our production research, the operating-layer controls paper, or explore DXAP, where DXRG research is applied. DXRG publishes this library and develops DXAP.
DXRG findings
- The Trading Opportunity List Is Part of the Agent’s Strategy
Our pre-alpha record isolates a selection effect at a rendered rank boundary. A candidate-list audit should precede claims about model market preferences.
- When Chosen Leverage Behaves Like a Configuration Constant
Historical agent leverage followed risk settings and explicit user numbers. Risk adaptation needs evaluation at the order boundary.
- Flat Leverage Across Volatility Sextiles Is a Risk Finding
Across 6,400 historical positions, median leverage stayed at 5x while volatility and liquidation rates changed sharply.
- A Profitable Price Excursion Is Different from a Profitable Trade
Historical positions often reached favorable prices and then closed negative. Capture analysis isolates an exit question from a direction claim.
- The Bracket Replay Identifies an Engineering Priority
A historical stop-target replay improved paired returns, while a trigger-quota failure exposed why protection needs execution verification.
- What We Do After a Directional-Edge Null
The historical fleet’s unfavorable return and signal results narrow the next experiment to measurable mechanisms and held-out prediction.
- A Paper Engine Needs Its Own Evidence Label
The historical fleet mixed mostly paper execution with a smaller real-capital book. Margin, funding and fill assumptions change what its results establish.
- Concrete Owner Instructions and the Historical Profit Cohort
Specific numeric instructions were associated with better outcomes in a bounded real-capital event. The next question is mandate design, with confounding preserved.
- Normalize Fee Eras Before Comparing Agent Returns
The historical fleet’s builder fees changed across three eras. Common-rate restatement separates cost assumptions from observed trading behavior.
- A Shared Harness Lineage Still Needs Transfer Tests
The two historical fleets shared configuration concepts but resolved text-setting conflicts differently. Transfer must name the component and environment.
- Simultaneous agent buys can emerge without a messaging channel
Read a concentrated historical buy event without converting shared timing into proof of agent coordination.
- One shared model can produce opposite token flows
Interpret the 92.9% trade share in two-sided windows using the correct denominator and time window.
- A sell-cascade count belongs to its definition
Keep the 3,878 historical cascades attached to the ten-vault, ten-minute rule before comparing populations.
- An inter-agent gap is not a causal response latency
Interpret the 9.5-second median gap in a large sell cascade without inferring a reaction clock.
- Capital withdrawn is different from trading profit
Separate the historical 92% withdrawal figure from return, residual value and cash-flow timing.
- Invocations to actions describe an operational funnel
Use the approximate 7.5 million invocations and 300,000 actions without treating inactivity as failure.
- Funded and active agents answer different population questions
Keep 3,505 funded vaults and 3,454 active vaults distinct when reporting participation.
- Token reaping made the tournament market state endogenous
Treat changing token survival as part of the historical market design rather than a background price feature.
- Rare failures can dominate an agent fleet loss
Read the historical liquidation count separately from its share of gross loss.
- Paper and live weights change what a fleet result means
Separate fills, positions, agents and market days before interpreting a mixed research fleet.
- A timestamp-artifact retraction changes the research contract
Use the withdrawn HIP-3 lag claim to specify what a point-in-time audit must preserve.
- A data cutoff is different from an analysis outcome window
Keep the August 15 freeze separate from the June 8 to July 26 historical P&L window.
- Aggregate inference tokens do not set a per-turn budget
Interpret roughly 70 billion tokens across 7.5 million invocations without inventing context or serving limits.
- Twenty-four prompt revisions create a selection history
Read repeated pre-launch repair as development evidence while preserving the trials behind a chosen prompt.
- Public paper artifacts support inspection at a declared level
Distinguish published aggregates, figures and extracts from a raw-data reproduction of historical results.
DXAP platform
- What a Persistent Trading Agent Must Remember Between Conversations
A worked lifecycle for persistent agentic trading: separate mandate, market state, and settled history before the next decision.
- How an Agent Wallet Fits a User-Owned Hyperliquid Account
Understand account ownership, delegated signing, and revocation through a concrete agent-wallet review scenario.
- Why Model Choice Matters More Inside a Shared Trading Harness
Compare trading models while holding the surrounding tools, policy, and account record fixed.
- How Social and Prediction-Market Signals Should Enter a Trading Turn
A provenance and timing framework for combining social signals, event odds, prices, and portfolio state.
- What Research and Chart Subagents Add to an Agentic Trading Turn
Separate research and chart tasks without letting worker agreement masquerade as independent evidence.
- Scheduled Trading Reviews and Agent-Triggered Revisits Solve Different Problems
Use periodic reviews for coverage and explicit triggers for changing conditions, with one reconciled decision state.
- What to Ask Your DXAP Agent After a Trading Decision
Review a trading decision through concrete questions about evidence, instructions, interpretation, and a proposed strategy revision.
- Why Trading Limits Need an Enforcement Point Outside the Model
A final-action test explains how deterministic policy differs from asking a model to respect a limit.
- Turn a Trading Conversation into a Strategy You Can Review
A revision workflow converts vague trading requests into explicit conditions while preserving the owner's intent and historical decisions.
- How Trading Research Should Change an Agent Platform
A research-to-product loop turns observed failures into fixtures, measured repairs, and separately verified platform improvements.
- Is your strategy ready for its first automated turn?
A prelaunch review that tests whether an owner can explain entries, waiting, and existing exposure.
- Which address should you inspect when an agent looks unfunded?
A diagnosis worksheet for account state and delegated signing authority.
- DXAP pricing and the cost of an executed round trip
A filled-volume worksheet separating builder fees, venue fees, funding, and price impact.
- How to read a public agent example with a paper label
An evidence review for strategy illustrations, public positions, and historical performance.
- When a chart setup needs an event-research question
A workflow for testing a catalyst thesis alongside completed-candle evidence.
- Choose a model with a latency and validity budget
A model review based on timely valid decisions rather than reputation alone.
- A supervision review for an agent that keeps running
A practical owner review covering active instructions, exposure, and completed decisions.
- A price trigger should reopen a question
How to distinguish a market wake-up from an authorized entry condition.
- Adapt a strategy to the instrument before reusing it
A portability review for symbols, leverage metadata, and exposure definitions.
- Review the source behind a social trading signal
An owner audit for origin, timestamps, repetition, and unsupported claims.
- Read the resolution rule before using prediction-market odds
A thesis review that distinguishes event probability from a perpetual-price forecast.
- Ask your agent what would invalidate its thesis
A conversation pattern for testing contradictions without demanding a trade.
- Diagnose a quiet agent before asking it to trade
A decision tree separating waiting, missing turns, rejected actions, and unfilled orders.
- Compare platform evidence instead of marketing superlatives
An architecture review that asks what a public claim actually establishes.
- Participating in alpha is an opportunity to give precise feedback
How to separate product feedback, observed trading outcomes, and public endorsements.
Frontier research
- Native probabilities versus verbal confidence in trading agents
A proposed calibration study separates token probabilities, spoken confidence, market forecasts and decisions under one saved evaluation set.
- Randomizing candidate lists to measure trading-agent selection bias
A proposed randomized render experiment tests how visibility and position affect asset selection independently of market information.
- Testing a volatility-conditioned position budget
A proposed sizing study tests whether volatility-aware exposure improves loss control under unchanged trade selection.
- An acceptance test for opening a protected position
A proposed failure-injection protocol defines when a trading position counts as protected and tests retries, quotas and recovery.
- Testing prediction-market signals with point-in-time provenance
A proposed experiment measures whether prediction-market inputs add information after timestamp, contract and liquidity checks.
- How to test parity between paper and live trading fills
A proposed execution study measures fill, funding and margin differences before using paper outcomes to assess an agent.
- Measuring the information value of research subagents
A proposed ablation distinguishes better information, changed decisions and net economic value when specialist research workers are added.
- Scheduled versus triggered revisits in an agentic trading loop
A proposed controlled experiment compares revisit timing while preserving information access, risk policy and agent exposure.
- Detecting trading regime changes without hindsight
A proposed online evaluation tests regime labels, detection delay and action value using only information available at each turn.
- A release regression corpus for an evolving trading harness
A proposed corpus tests mandate, input, execution and recovery behavior so improvements can be evaluated across DXAP harness releases.
- Can one trading intent survive a change of venue?
A proposed adapter experiment tests economic meaning across synthetic venue rules.
- Does a chart image add information beyond its numbers?
A proposed paired study separates visual representation from additional market information.
- How much does delayed news change an agent decision?
A randomized delay experiment measures the effect of information arrival on fixed decision opportunities.
- Which opportunities deserve a limited research budget?
A proposed allocation study measures evaluated opportunity coverage under a fixed total budget.
- What should a trading runtime process first during a trigger storm?
A proposed queue experiment prioritizes exposure management under bursts of agent wakeups.
- When human help changes an agent evaluation
A proposed protocol records owner interventions and separates assisted outcomes from autonomous behavior.
- How to validate an execution-impact model before relying on it
A proposed holdout study tests impact estimates with matched order-size and market-state controls.
- An event probability is only one part of a price forecast
A proposed two-stage forecast separates event occurrence from conditional market response.
- How heterogeneous agents coordinate through a shared market
A proposed shared-state experiment distinguishes correlated reaction, communication and price feedback.
- Planning when a tool answer may arrive incomplete
A proposed planning study varies tool reliability and tests how agents value information before acting.
- What an execution rehearsal can prove without trading authority
A proposed rehearsal protocol tests the complete planning path while withholding submission capability.
- Trading-hour transitions need explicit agent fixtures
A proposed synthetic calendar suite tests decisions around open, close and scheduled market interruptions.
- A good trade score can still produce a poor portfolio
A proposed objective comparison evaluates interactions among simultaneous positions.
- Register an agent-release evaluation before its outcomes arrive
A proposed release protocol freezes questions, comparisons and decision rules prospectively.
- When should a failed trading-agent theory be retired?
A proposed research decision record distinguishes a tested failure from an unresolved or untested mechanism.
Learning theories
- Trading Agent Feedback Labels: Which Outcome Should Teach the Agent?
A proposed label design separates mandate compliance, execution quality and market outcomes before trading traces enter an agent evaluation.
- Outcome Bias in Trading Agent Grading: A Blinded Test
A paired review experiment checks whether a profitable result changes how reviewers judge the same trading-agent decision.
- Process Reward vs Economic Reward for Trading Agents
A proposed two-axis evaluation tests whether better trading-agent process scores translate into decision economics under realistic costs.
- On-Policy Trading Data: Track Which Agent Produced Each Example
A provenance design for distinguishing current-policy trading trajectories from historical traces before evaluating a learning proposal.
- Probability Calibration and Trading Policy: Test Them Separately
A worked example shows how a calibrated forecast can produce different trading actions under different cost and risk policies.
- Separating Prompt Order from Context Length in Trading-Agent Ablations
A four-arm trading-agent experiment crosses assigned context length with instruction placement to distinguish salience changes from action improvements.
- Model Updates vs Harness Updates: Attribute Trading-Agent Improvements
A proposed component comparison distinguishes a stronger trading model from improvements to instructions, state or tools.
- Distribution Shift in Trading Agent Learning: Which Change Broke Transfer?
A proposed shift audit separates changing markets, user mandates and agent infrastructure before judging learning transfer.
- Counterfactual Trading Replay: Check Action Support Before Comparing Policies
A proposed replay audit identifies which alternative trading actions the saved data and simulator can actually evaluate.
- Cost-Aware Evaluation for Trading Research Subagents
A proposed marginal-value experiment tests when research or chart subagents improve trading decisions enough to justify their cost and delay.
- Audit the Teacher Before Distilling Trading Traces
Trace imitation can transfer a teacher’s unsupported rules along with its useful decisions.
- Missing Action Propensities Break an Off-Policy Ratio
A saved action and token log probabilities may still leave the behavior-policy denominator unknown.
- Test Representation While Holding Information Fixed
A representation test needs a fact-level equivalence check before model scores are compared.
- Reward Normalization Across Market Days Changes the Lesson
Per-day standardization can erase an economically important difference between quiet and volatile sessions.
- Assigning Delayed Credit to Agent Actions
A later portfolio outcome needs a declared path back to the decisions that could have affected it.
- Contradictory Feedback Can Make an Agent Oscillate
Before changing the learner, check whether opposite labels refer to the same mandate and information.
- Separate Model Uncertainty from Market-Data Uncertainty
More model agreement cannot repair a stale or ambiguous observation.
- Ensemble Agreement Can Hide Shared-Source Dependence
Several models can repeat one source error and still look unanimous.
- Weight Agent Errors by Operational Severity Carefully
Severity weights encode an objective and should remain visible beside ordinary error counts.
- Test Whether New Learning Forgets Mandate Compliance
A candidate can improve on its new task while losing an older instruction boundary.
- Choose Verifiers That Expose Reward Gaming
A verifier should discriminate substantive compliance from text that merely resembles a passing answer.
- Checkpoint Selection Can Leak the Evaluation Set
Repeatedly choosing checkpoints on a holdout turns that holdout into a development signal.
- Keep Execution Outcomes at the Right Training-Label Boundary
A failed venue submission should teach the stage responsible for the failure.
- Evaluate Learned Tool Routing Against a Fixed Policy
A learned router earns its complexity by choosing useful information under the same available tools.
- Evaluate a Revised Rationale While Holding Actions Fixed
Clearer explanations deserve their own evidence rather than a trading-performance claim.
Mandates and reasoning
- When a Trading Slider and Strategy Text Disagree
A worked conflict case shows how authenticated controls and strategy text should resolve into one inspectable trading mandate.
- A Number in a Trading Mandate Needs a Unit and a Trigger
Why a concrete trading instruction needs quantity definitions, conditions and expiry before its behavior can be evaluated.
- Temporary Trading Instructions Need an Expiry Event
A worked news-window example separates expiry of permission, pending execution and ongoing position management.
- Which Strategy Version Authorized This Trade?
A race between a strategy update and an unfinished turn shows why every action needs one attributable mandate version.
- An Observation Is Evidence; an Action Rationale Is a Decision Record
Separating market facts, interpretation and rationale makes a trading trace easier to evaluate and prevents earlier reasoning becoming policy.
- What a Fee-Sentence Reorder Actually Demonstrated
DXRG’s fee-placement result shows a large change in trace citation under fixed facts, with clear limits on economic interpretation.
- When an Agent Invents a Trading Rule from Its Own History
DXRG’s rule-fabrication finding supports explicit authority labels and compound-intervention attribution, with a targeted regression design.
- Keep Market Facts Separate from Trading Instructions
A news excerpt can describe someone else’s imperative without authorizing an agent to act; source roles make the difference inspectable.
- A Valid Trading Action Needs More Than Valid JSON
A quantity conversion example shows how schema checks, mandate policy and final-payload validation answer different questions.
- No Trade Is a Decision Worth Recording
An explicit abstention record separates a thesis-driven wait, insufficient data and failed execution in trading-agent evaluation.
- Gross and net exposure limits measure different risks
Separate total absolute position notional from signed directional exposure with a two-position mandate fixture.
- Define the high-water mark before enforcing drawdown
A cash-flow-aware peak-equity fixture makes a drawdown restriction reproducible.
- Daily loss limits need an explicit reset clock
Use a UTC interval fixture to distinguish loss-budget boundaries from the owner’s displayed calendar day.
- Resolve strategy priorities with a decision table
An illustrative trend-versus-reversion conflict makes precedence visible before execution.
- Conflicting risk bands need a feasible intersection
An interval fixture distinguishes a valid shared budget from a mandate that requires owner resolution.
- Research workers provide evidence; owners grant authority
A forged amendment fixture separates a specialist’s information from permission to change a mandate.
- Review strategy settings as a field-level diff
A three-field proposal fixture reveals changes that a conversational summary can hide.
- Review simultaneous strategy proposals against one base
A conflicting two-proposal fixture prevents accidental combination of separately reviewed changes.
- Allowlist amendments need a rule for existing positions
An ETH-removal fixture separates entry eligibility from authority to manage an outstanding holding.
- Minimum data requirements belong in the mandate contract
A three-snapshot fixture defines when a strategy has enough information to propose an entry.
- Define whether pause stops entries or position management
A held-position fixture turns an ambiguous pause command into explicit permitted actions.
- Separate enforceable controls from advisory preferences
A ceiling-and-patience fixture distinguishes a machine-checkable restriction from guidance for model judgment.
- A slippage budget needs a reference price and side
A buy-and-sell price fixture compiles execution tolerance without hiding fees inside the same number.
- Aggregate related instruments with an explicit risk map
A spot-and-perpetual fixture separates directional offsets from gross family exposure.
- Validate compiled strategy units with invariant tests
A generated-case protocol checks authorization properties across many typed strategy inputs.
State and memory
- Portfolio Reconciliation Before an Agent Adds Exposure
A proposed fixture for delayed fills, fees, and the portfolio version used to approve the next trading action.
- Superseded User Instructions in Trading-Agent Memory
How to test a strategy update when old instructions remain easy to retrieve and a model call is already running.
- Memory Provenance Versus Trading Authority
A source can be correctly identified while remaining unable to authorize a trade or change an agent mandate.
- When News Should Expire From a Trading Agent’s Context
A dependency-based fixture for headlines that remain retrievable after the event they describe has resolved.
- Handling Tool Results That Arrive After an Agent’s Decision
A proposed procedure for late research responses whose original request context has already been superseded.
- Calendar Time Versus Turn Count in Agentic Trading
A proposed schedule-invariance test for agents that confuse repeated observations with elapsed market time.
- Testing Whether Retrieved Memory Is Relevant to a Trade
A relevance fixture that separates semantic similarity from current decision eligibility and downstream usefulness.
- External State Versus Model Belief in a Trading Runtime
How to test a model that insists an order filled when the authoritative venue state says it remains open.
- Preventing Cross-Agent Context Contamination in Trading Research
A proposed test for subagent results that accidentally import another account’s limits, positions, or mandate.
- Compact State Snapshots Versus Long Context for Trading Agents
A proposed representation comparison that preserves modern context capacity while testing what current decisions actually need.
- Event Logs and Mutable Portfolio Snapshots
Choose the replay boundary before choosing how to store portfolio state.
- Bitemporal Corrections in Agent State
Preserve both corrected history and the information available at the original decision.
- Causal Ordering Between Account Events
Account state needs explicit dependencies when timestamps cannot order its changes.
- Memory Tombstones and Deletion Lineage
Deletion needs to cover derived memories as well as the original item.
- Confidence Labels for Unverified Source Facts
Separate a fact’s verification status from the model’s confidence in its interpretation.
- Resolving Contradictory Source Facts
Resolve a source conflict at the claim and scope level before selecting a current fact.
- Auditing Information Loss in Agent Summaries
Measure the decision-relevant facts a summary drops, changes, or merges.
- Retrieval Recall Under a Fixed Token Budget
Compare memory retrieval by required evidence delivered inside the same rendered budget.
- Numerical Precision During Context Compression
Retain exact operational numbers and compress the surrounding explanation.
- State Digests Across Agent Tools
Use digests to compare declared state representations, with schema and semantics attached.
- Invalidating Agent State Caches After a Trade
A trade transition should invalidate every derived view that depends on its effects.
- Clock Skew Sensitivity in Agent State
Test decisions across plausible clock offsets instead of treating timestamp comparisons as exact.
- Quarantining Scenario Memory in Agent Evaluations
Keep counterfactual and stress-case memories from becoming factual history.
- Memory Retention and Operational Observability
Choose retention by the incidents you must reconstruct and the information each record exposes.
- Human and Machine Readable State Parity
A readable state view should preserve the operational meaning of its typed source.
Trace evaluation
- A Minimum Decision Trace Schema for Trading Agents
A proposed schema that joins agent intent, policy checks, execution and reconciliation without treating a rationale as an execution receipt.
- Freezing Replay Inputs Before Comparing Trading Models
How to define the replay boundary, preserve captured context and avoid silently evaluating a different information set.
- Tool Availability Is Part of a Fair Trading-Model Comparison
A procedure for separating tool-use skill from information access in trading-agent model evaluations.
- Counting Failed and Missing Turns in Agent Evaluations
Why completed-turn scores need scheduled-turn denominators and explicit missing-status reporting.
- What a Trading-Agent Rationale Can Actually Establish
How to test explanation consistency against recorded inputs and actions while keeping causal claims bounded.
- Evaluating Trading-Agent Abstention Without Rewarding Silence
A coverage-aware method for comparing no-trade behavior with conditional decision quality.
- Choosing Deterministic and Probabilistic Graders for Agent Traces
Separate executable constraints from semantic interpretation and measure disagreement before using an LLM grader.
- Attributing a Trading-Harness Improvement to Its Actual Intervention
How to describe compound fixes, isolate components and retain the historical scope of measured trace improvements.
- Separating Execution Errors from Prediction Errors in Agent Trading
A stage-by-stage diagnosis of trading losses that distinguishes forecasts, sizing, policy and venue execution.
- Recording Model and Harness Versions for Trading-Agent Results
Why a model label alone cannot identify the configuration that produced a trading decision or evaluation score.
- Shared Examples Can Contaminate an Agent Benchmark
A case-family audit separates copied examples, semantic siblings and genuinely new trading decisions.
- Model Alias Drift Can Break Replay Reproducibility
Resolve model identity at request time and quarantine ambiguous alias changes before comparing agent behavior.
- Detect Prompt Truncation Before Scoring Agent Decisions
A section receipt and boundary-marker fixture identify missing instructions before an agent trace receives a behavioral grade.
- Classify Malformed Tool Arguments by the Repair They Need
A staged taxonomy distinguishes serialization, schema, domain and authority failures in trading-agent tool calls.
- Timeout Censoring Changes What an Agent Evaluation Can Estimate
Distinguish deadline failure from an unobserved completion time and preserve timed-out cases in quality reporting.
- Latency Quantiles Reveal Trading-Agent Tail Behavior
Compare latency distributions and deadline misses rather than relying on average response speed.
- Choose a Completion Stage Before Setting an Agent Service Objective
Define separate objectives for usable proposals, submitted actions and reconciled outcomes.
- Deduplicate Traces by Economic Intent Without Losing Retries
A three-identity audit prevents transport retries from becoming additional trading decisions.
- Stratify Manual Audits Around the Failures You Need to Understand
A weighted review design finds rare severe traces while preserving a population estimate.
- Reviewer Agreement Depends on Failure Prevalence
Report the agreement table alongside kappa so rare defects remain visible.
- Error Severity and Failure Frequency Need Separate Reports
A severity-aware trace audit avoids treating common harmless defects as equivalent to rare unauthorized actions.
- Attribute Tool-Chain Failures at the Broken Dependency
A dependency audit separates the first corrupted result from later calls that merely propagate it.
- Tool-Output Schema Migrations Need Consumer Tests
A compatibility matrix checks old and new consumers against changed trading tool responses.
- Separate Replay Determinism from Model Choice Variation
A two-stage replay fixture distinguishes runtime instability from stochastic changes in model proposals.
- Unsuccessful Tool Calls Are Information-Access Failures
Evaluate how agents respond when needed evidence is unavailable instead of grading only the final narrative.
Forecast evaluation
- Calculating Brier Score for an Agent Forecast Audit
A five-row worked example for checking probability scores, sample alignment, and comparison baselines.
- When Log Loss Exposes an Overconfident LLM Forecast
An illustrative comparison shows how a confident miss can outweigh several confident hits.
- Choosing Bins for an LLM Forecast Reliability Diagram
A worked example shows why bin counts and boundaries belong beside a calibration chart.
- Separating Calibration from Discrimination in Agent Forecasts
Two constructed forecasters show why probability meaning and event ranking require separate checks.
- Building a Time-Valid Base-Rate Baseline for Agent Forecasts
A fixed reference probability prevents an impressive accuracy number from becoming a misleading skill claim.
- Purging Label Intervals at an Agent Evaluation Boundary
A timestamp example explains why chronological rows alone do not prevent information leakage.
- A Walk-Forward Protocol for Native LLM Forecast Comparisons
A three-window example separates model updating, probability calibration, and honest aggregate reporting.
- Why Overlapping Forecast Labels Are Not Independent Evidence
A rolling-horizon fixture shows how many saved rows can represent far fewer distinct market intervals.
- Recording the Trials Behind a Winning Agent Forecast Score
A simple chance calculation explains why a selected winner needs an untouched evaluation.
- Confidence Intervals When Agent Outcomes Are Sparse
A zero-event example shows why a small observed rate can still leave substantial uncertainty.
- Target Prevalence by Asset in LLM Forecast Evaluation
Measure event frequency separately for each asset before interpreting a pooled forecast score.
- Calibration Slope and Intercept for a Native LLM Forecast Head
Use a logistic diagnostic to separate shifted probabilities from excessive confidence.
- Normalizing Multiclass Native LLM Forecasts
Make candidate probabilities coherent while preserving the mass outside the candidate set.
- Probability Rounding and Reproducible Forecast Scores
Keep scoring precision separate from the probability shown to a reader.
- Paired Forecast Score Differences Versus Separate Means
Compare identical questions row by row before estimating uncertainty.
- Risk and Coverage for Selective Agent Forecasts
Evaluate confidence-based forecast selection across retained fractions.
- Conformal Forecast Coverage Under Temporal Shift
Audit empirical coverage and set size when market distributions change.
- Auditing Disagreement in Forecast Resolution
Separate price-source ambiguity from native LLM probability error.
- Precision and Recall for Rare Market Event Forecasts
Use event counts to expose false alarms hidden by overall accuracy.
- Cost-Sensitive False Positives in Agent Decision Screening
Evaluate an internal review screen using declared asymmetric error costs.
- Macro and Micro Averaging Across Trading Agents
Choose whether the research question concerns a typical agent or a typical forecast.
- Label Noise and the Forecast Benchmark Ceiling
Quantify the score floor implied by an explicit observation-noise model.
- Censored Outcomes in an Agent Forecast Evaluation
Keep unresolved future outcomes visible in the denominator.
- Scoring Native LLM Forecasts by Volatility
Use lagged volatility strata to diagnose where probability error changes.
- How Forecast Scores Depend on the Resolution Horizon
Treat each horizon as a distinct target before comparing model quality.
Execution mechanics
- Maker and taker fees: a fill-level arithmetic check
Calculate mixed maker and taker fees without mistaking an order type for an execution classification.
- Spread crossing cost for an agent execution
Measure the bid-ask component of execution cost with a declared reference price and quantity.
- Estimating execution VWAP from a saved order book
Walk book levels to calculate an illustrative execution estimate while preserving its snapshot limitations.
- Partial fills: reconstruct quantity before judging completion
A worked partial-fill ledger for autonomous agents, including duplicates, residual quantity, and cancellation.
- Cancel and replace races in autonomous order handling
Trace a replacement instruction through a late fill so residual exposure remains explicit.
- Order retries after a timeout: preserve one economic intent
A recovery procedure that separates transport uncertainty from a new trading decision.
- An order acknowledgement is not a fill receipt
Keep accepted requests, live orders, executions, and final reconciliation distinct in agent traces.
- Reduce-only orders and the agent position they reference
Check closing intent against position mode, live quantity, and venue-specific reduce-only semantics.
- Tick and lot rounding as an agent authorization boundary
Normalize executable prices and quantities without silently widening an agent instruction.
- Funding cash flows: reconcile settlements separately from trade profit
A signed funding ledger that distinguishes displayed rates, estimated charges, and settled account entries.
- IOC, FOK and GTC: Preserve the Agent’s Execution Deadline
Time-in-force is part of trading authorization. A worked liquidity case separates immediate execution, resting exposure and unsupported FOK intent.
- Post-Only Rejection Needs a New Price Decision
A moving quote can invalidate a maker-only order. Preserve the liquidity instruction and measure rejection handling separately from fill quality.
- Market Execution Needs an Explicit Worst Price
Translate an urgent trading instruction into a signed price boundary. A worked buy case distinguishes execution urgency from unlimited spending.
- Order Amendments Need a Queue-Priority Assumption
A modification can change waiting economics. Record amendment lineage and treat priority retention as venue-specific behavior requiring explicit evidence.
- Execution Price Improvement Requires a Frozen Reference
Measure favorable execution against a declared pretrade reference. A worked ledger separates price improvement, fees and subsequent market movement.
- Terminal Order Statuses Need Reason-Preserving Mapping
A canceled order can reflect owner action, margin changes or venue restrictions. Preserve reasons and reconcile economic effects before closing the agent record.
- Cancel All Orders Is Different from Clearing an Account
After a broad cancellation, verify orders, positions and collateral separately. A worked account case defines a narrower, reviewable stopping condition.
- Concurrent Agent Actions Need One Nonce Owner per Signer
Nonce tracking follows the signing wallet. A concurrent-worker fixture defines atomic allocation and separates replay protection from execution sequencing.
- Websocket Reconnects Need an Execution Backfill Boundary
Reopening a stream restores transport. A gap fixture combines historical fills, stable economic identity and a coverage-aware account reconciliation.
- Self-Trade Prevention Can Remove an Agent’s Resting Quote
Two strategies on one address can interact without a trade receipt. A worked book case tracks the resting-order cancellation and the surviving aggressive intent.
- Cross and Isolated Margin Need Different State Snapshots
Collateral sharing changes the meaning of available capacity. A simplified two-position case defines account abstraction and margin-mode fields for agent decisions.
- Execution Fees Need a Currency and a Valuation Time
A fee ledger should preserve native amounts before conversion. A two-rate example separates execution cost from later changes in the fee asset’s price.
- Outstanding Orders Must Reserve Potential Agent Exposure
An unfilled order can consume a future exposure budget. A simultaneous-worker example defines worst-case reservations without assuming offsetting orders execute together.
- Realized and Unrealized P&L Need Separate Agent Labels
A partial close changes both accounting categories. A simplified position case separates closed trade profit, remaining marks and independent cash movements.
- Onchain Execution Needs an Explicit Finality Contract
A chain identifier and an API observation establish different facts. Define the event evidence required before dependent agent actions and cross-system reports advance.
Market data
- Event Time vs Receipt Time in Agentic Trading Data
A worked timestamp example for reconstructing what a market-data consumer knew at a decision boundary.
- As-of Joins for Trading Agents: Enforce Information Availability
A small quote-joining example that exposes future matches, stale matches, and ambiguous ties.
- Candle Close Leakage in Agentic Trading Replays
A minute-bar example for separating bucket labels from feature availability.
- Missing Market Candles: A Data Contract for Trading Agents
A gap-classification workflow that keeps no-trade periods separate from collection failures.
- Token Decimal Precision for Onchain Trading Agents
A token-decimals example with exact arithmetic, metadata provenance, and failure handling.
- Asset and Instrument Identity in Agentic Trading Data
How to build instrument keys that distinguish venues, pairs, chains, and metadata changes.
- Order-Book Depth Units for Autonomous Trading Agents
A three-level example showing how depth differs from a best quote and why aggregation matters.
- Crossed Order Books: Diagnosing Agent Market Inputs
A reconstruction checklist for negative spreads, mismatched snapshots, and venue state.
- Stale Quotes and Feed Health in an Agent Trading Loop
A quiet-market example for separating old prices from broken subscriptions.
- Event Sampling Bias in Trading-Agent Market Features
A two-regime example showing how feed activity changes an average even without a price change.
- Reconstruct the Instrument Universe an Agent Could Actually Trade
Build an eligibility ledger that preserves listings, restrictions, and removals before comparing agent selection.
- Keep Delisted Markets in an Agent Research Denominator
A four-market fixture shows how a current catalogue can reverse a historical market-selection summary.
- Recover a Sequence Gap Before Reusing an Agent Order Book
Use a feed-specific continuity check and a recovery receipt to identify when reconstructed market state becomes usable again.
- Align an Order-Book Snapshot with Buffered Updates
A sequence-boundary fixture and absolute-size example catch two common errors in reconstructed depth.
- Separate Funding Estimates from Settled Funding Records
An illustrative rate change shows why agent features and account funding cash flows need different data records.
- Choose the Price Definition Before Calculating an Agent Feature
A three-price fixture shows how an unnamed price changes portfolio valuation and market interpretation.
- Keep the Revision History of News and Economic Events
A revised-release fixture preserves the earlier agent input while allowing a later corrected account of the event.
- Version the Resolution Wording Behind a Prediction-Market Signal
A synthetic deadline amendment shows why a market probability needs the question and rules version that gave it meaning.
- Count Reposts as Distribution, Separate from Independent Claims
A post-reference fixture keeps social reach visible while preventing copied claims from becoming false corroboration.
- Preserve Source Disagreement About an Event Time
A release-time conflict fixture treats disagreement as data and tests whether it can change feature admission.
- Preserve Hyperliquid Asset IDs Through Numeric Conversion
An integer-width fixture and metadata binding check prevent valid JSON from routing an action to the wrong asset.
- Compare REST and WebSocket Observations by Semantic State
A versioned quote fixture separates normal observation lag from a true adapter disagreement.
- Retain the Market Inputs Needed to Reconstruct an Agent Decision
A deletion fixture shows why prompt summaries and hashes cannot substitute for missing source payloads.
- Handle Duplicate Records at Historical API Page Boundaries
An inclusive-boundary fixture preserves equal-time events while preventing repeated rows from inflating a market feature.
- Convert Account Value with the Rate Available to the Decision
A two-rate fixture separates native-currency performance from reporting-currency movement.