Define an active review before measuring engagement

By DX Research Group · · Incentives and participation

A semantic review contract separates inspected decisions from open tabs, visualizer time and autonomous agent operation.

An agent can run while its owner is away, and a browser can stay open while nobody is reading. Measuring either as owner supervision confuses autonomous runtime with human attention. We would define an active review through what the owner inspected and understood.

The distinction is concrete in DXAP. The Chapter 0 guide says the visualizer can keep depicting a journey between turns and that closing it leaves the agent running. The activity guide instead directs owners to inspect a turn and compare its outcome with positions and trades. Those are different behaviors with different evidence value.

A review has an object and an answer

For a proposed 14-day measurement study, define a qualifying review as opening a specific recorded turn, inspecting its decision state and correctly answering one sampled question about that state. The question might ask whether a submitted order has a recorded fill, or whether a completed turn chose to wait. A correct answer shows a narrow act of interpretation. It provides stronger evidence than elapsed tab time while remaining an incomplete measure of full understanding.

Invite 80 consenting owners and offer the same review opportunity. Track whether each owner completes at least one qualifying review during the window, reviews per owner and distinct turns reviewed. Keep all 80 in the participation denominator. Sample questions sparingly so measurement effort does not become the main reason people visit.

An illustrative owner leaves the visualizer open for 120 minutes, spends four minutes in Activity and correctly interprets two turns. The record contains 120 visualizer minutes and two qualifying reviews. Another owner keeps Activity visible for 45 minutes while away and submits no interpretation. Their measured review count is zero, with understanding unobserved. It would be excessive to conclude that they understood nothing; the study only lacks affirmative review evidence.

Now suppose 50 of 80 owners have ten minutes of open-tab time, while 24 of 80 complete a qualifying review. The two participation rates are 62.5% and 30%. Their gap identifies how much the looser proxy can inflate a claim about measured supervision. These are illustrative counts, not product analytics.

Validate the definition against real reading

Use a consenting subsample for brief observed sessions. Compare the contract with a reviewer assessment of whether the owner actually interpreted a decision. Record false positives from guessing and false negatives where thoughtful reading occurred without an answer. Rotate question variants and avoid requiring a trade, instruction change or approval to count as review.

No-trade turns qualify equally when the owner correctly identifies the missing conditions. Multiple refreshes of the same turn count once for coverage, with later reconsideration recorded separately if new evidence arrives. This makes the measure resistant to repeated clicking while allowing useful follow-up.

The resulting report can say how many owners demonstrated a defined review behavior in 14 days. That is a bounded claim about supervision. It gives us a better target for testing explanations and feedback tasks than tab duration, and leaves forecasting quality and capital outcomes to their own measurements.

Sources

Related field notes