When human help changes an agent evaluation

By DX Research Group · · Frontier research

A proposed protocol records owner interventions and separates assisted outcomes from autonomous behavior.

Human intervention changes the treatment being evaluated. We propose recording every owner review, edited instruction and emergency action, then separating assisted operation from autonomous operation. The question is how the intervention changes decisions and outcomes within a declared population. This is an unrun protocol, with no new measured result.

The continuous record companion reports historical differences among owner interaction cohorts. Those cohorts can reveal association while leaving selection unresolved: owners who intervene may differ in expertise, risk tolerance or the difficulty of their agents' situations. The controls paper companion describes owner controls and authenticated mandates, providing a place to attach intervention events to the trace.

A rescue is part of the outcome

Consider an illustrative sample of 100 evaluation episodes. Owners intervene in twenty. Sixteen of those twenty end with a policy-valid final state; sixty-four of the remaining eighty do too. The overall rate is 80%. Both observed groups also show 80%, but the matching percentages establish no causal equivalence. Humans might have rescued difficult cases while selecting which cases to review. An episode that required a rescue differs from an episode the agent completed unaided.

Create an intervention ledger with the actor, time, instruction version and affected action. Separate a comment that adds context from an edit that changes a risk limit or a manual position close. Record the state immediately before the intervention, including any unresolved order. An owner cancellation after a bad proposal belongs in the episode history even when the cancelled trade never settles.

A first descriptive report should show completion with and without intervention, time to assistance and the reasons assistance was requested. It should keep the original denominator of assigned episodes. Excluding rescued cases from one group or deleting cancelled proposals would make the system appear more autonomous than it was.

Assign review availability before the case unfolds

Where an offline study can safely randomize assistance, assign access to a reviewer before the market outcome is known. Give reviewers a fixed rubric and identical saved context. Compare assigned availability as the primary treatment, then report actual intervention as a secondary behavioral measure. This avoids estimating the effect only among cases where someone decided to help.

We would also score whether the agent recognizes that authority changed, refreshes its state and attributes the next decision to the revised mandate. Human assistance can improve an outcome while exposing an attribution failure. Reviewer time belongs in the cost account, including inspection that produces no edit.

The resulting report would describe the whole human-agent configuration: the agent alone, review availability and observed assistance. A useful conclusion could be that periodic review improves constraint compliance in particular difficult cases. That conclusion would still require its assigned-review evidence and would remain distinct from a claim that the unaided agent improved.

Sources

Related field notes