Calendar Time Versus Turn Count in Agentic Trading
By DX Research Group · · State and memory
A proposed schedule-invariance test for agents that confuse repeated observations with elapsed market time.
Ten agent turns can happen in ten seconds or ten hours. A cooldown, news window, or settlement wait should use the clock relevant to its rule. We would test whether changing invocation frequency alters a time-dependent decision when the underlying event timeline stays fixed.
Our state and memory research describes a historical cadence failure: some agents treated elapsed ticks since the prior trade as a signal. Filtering repeated observations reduced that pattern in sampled traces, with no numeric incidence rate reported.
Hold the market timeline still
Construct an illustrative event record with a trade at 12:00 and a mandate requiring a 30-minute cooldown before additional exposure. Replay the same quotes and portfolio events under two schedules. One invokes the model every minute; the other invokes it every ten minutes.
At 12:20, both schedules remain inside the cooldown. The first has seen many more prompts, but elapsed wall time is identical. At 12:30, both reach the eligible boundary if the rule uses that clock and all other conditions remain satisfied.
Give each snapshot a current time, last relevant event time, and derived elapsed duration. Preserve timezone and the comparison operator at the boundary. “After 30 minutes” and “at least 30 minutes” can have different interpretations unless the compiler makes the rule explicit.
Turn count remains useful for diagnosing runtime behavior. It can identify repeated requests, retry loops, or expensive invocations. Its useful diagnostic role does not grant it market-time meaning.
Test duplication as well as frequency
Add several duplicate observations between 12:10 and 12:11. The market event sequence remains unchanged while context grows. A model may interpret the repetition as prolonged inactivity or increasing evidence. The fixture asks whether that change affects the typed action or its rationale.
We would score premature actions, boundary errors, and invented cadence rules separately. Also record valid no-trade decisions. A response that declines the trade because of the active cooldown has followed the intended temporal rule even if the model has a positive market forecast.
The published evaluation registry freezes model, sampling, mandate, and policy components for state-transition tests. Its fields remain unrun. The schedule-invariance experiment here adds paired invocation calendars and duplicate observations as explicit interventions.
Interpret the historical observation precisely
The operating-layer paper companion places cadence drift among the historical harness failures. The qualitative reduction supports investigating repeated observation effects. It supplies neither an isolated causal estimate for a particular clock field nor a current DXAP pass rate.
DXAP publicly describes scheduled turns and triggers. A useful acceptance test for that architecture checks that changing the trigger frequency preserves calendar-based constraints. This would establish a scheduling property within the tested cases, while performance comparisons still need a separate market evaluation.
We would publish the event timeline and both invocation schedules beside the result. That lets another team reproduce the exact 12:20 and 12:30 decisions. Continue with mandate compilation for translating human time language into an explicit executable rule. The strongest temporal record connects the intended clock, its timestamps, the computed duration, and the final policy outcome.