Clock Skew Sensitivity in Agent State
By DX Research Group · · State and memory
Test decisions across plausible clock offsets instead of treating timestamp comparisons as exact.
State freshness should account for uncertainty in the clocks being compared. We would test a decision across plausible clock offsets and identify the cases whose admission changes. A single timestamp difference can hide a substantial error when a remote source and the runtime have different clock estimates.
The clock fields in our state and memory framework establish the inputs. Our trace feedback method supplies the record for comparing admission and resulting action. This note concerns offset sensitivity, rather than the general distinction between event time and arrival time.
Age becomes an interval
Suppose a hypothetical source labels a quote 12:00:00.000 and the runtime receives it at local time 12:00:00.800. The apparent age is 800 milliseconds. If the source clock is 500 milliseconds ahead of the runtime's time basis, the actual age on that basis is 1,300 milliseconds. If it is 500 milliseconds behind, age is 300 milliseconds.
A one-second freshness threshold therefore lies inside the possible age interval, 300 to 1,300 milliseconds. We would classify that quote as boundary-uncertain under this stated offset range. The example supplies no evidence that any actual feed has that clock error; it is a test input selected to expose sensitivity.
A local receipt-age threshold avoids this particular cross-clock subtraction, but it answers how long the runtime has held the message. It cannot by itself bound how old the observation was before delivery. Both measurements can be useful if their meaning remains explicit.
Sweep offsets, hold the facts fixed
Our proposed fixture uses the same quote, portfolio, and mandate while changing only the assumed source offset. It records whether the assembler admits the quote and whether the model's proposal changes after admission. That separates a deterministic freshness-boundary effect from a model response to different visible state.
We would include offsets of minus 500, zero, and plus 500 milliseconds, plus a quote far from the threshold. The latter should remain stable across the sweep. If every case changes, the fixture or implementation may be coupling offset metadata to unrelated behavior. Unknown offset should have its own declared response rather than inheriting zero.
Lamport's treatment of distributed clocks explains why event order and physical time are different tools. A logical order can preserve dependencies without providing a physical-age estimate. Conversely, a physical clock estimate can support age bounds while leaving unrelated events unordered.
The output should name the smallest tested offset that changes admission and the population of boundary-sensitive cases. Those are sensitivity measurements, with no implied estimate of production error frequency. A team can then decide whether tighter clock measurement or a wider operational margin addresses the observed weakness. The practical result is a freshness rule that states its clock assumptions and exposes uncertainty where those assumptions control the next action.