Contradictory Feedback Can Make an Agent Oscillate

By DX Research Group · · Learning theories

Before changing the learner, check whether opposite labels refer to the same mandate and information.

A policy can oscillate because its feedback contains incompatible objectives, rather than because the optimizer is simply unstable. We would resolve the conditioning variables behind contradictory labels before reducing a learning rate or collecting more examples.

Learning from human preferences provides a precedent for using comparisons as supervision. In a trading setting, the comparison must preserve owner authority and state. Our operating-layer controls connect action validity to the mandate, and trace feedback retains the context needed to explain disagreement. The following audit is proposed.

An apparent contradiction may be a missing condition

Take an illustrative state where a position has gained 5%. Evaluator A rewards holding because the mandate asks for a long horizon. Evaluator B rewards closing because another mandate asks for realization at 5%. If the dataset strips the mandate and retains only market state, the same visible input receives opposing action targets.

For a binary hold-versus-close learner, ten hold labels and ten close labels on the same input imply an empirical close frequency of 0.5. The minimum cross-entropy prediction is therefore 0.5. Training first on the hold batch and then on the close batch can move the prediction back and forth, especially with large updates. The labels have omitted the variable that would make the task learnable.

Restoring the mandate partitions the cases into two distinct conditional problems. The audit should also check label version, user horizon, and whether the evaluator saw future outcomes. Two judgments on identical visible inputs may still be legitimate if their objectives differ, but the learner needs those objectives represented.

Diagnose disagreement before averaging it away

We would group apparently contradictory examples by exact decision-input identity, then inspect the evaluator instruction and permitted evidence. A disagreement matrix distinguishes factual mistakes, conflicting objectives, ambiguous mandates, and ordinary evaluator uncertainty. Each group receives a different disposition.

Factual mistakes can be corrected with source evidence. Conflicting objectives require separate targets or explicit conditioning. Ambiguity may warrant a soft label or exclusion from a hard action target. Averaging everything into a majority label hides these mechanisms and can erase a valid minority mandate.

The proposed evaluation holds out complete mandate families rather than scattering near-duplicate turns across splits. It tests whether the candidate obeys each mandate in the same market state and whether a recent batch causes regression on earlier mandates. Inspect action probabilities and actual typed actions separately because small probability changes can cross a decision threshold.

The practical stopping point for this audit is a reconciliation table with the restored condition and a supported target for each conflicting pair. If conflicts remain after all relevant context is supplied, we would report evaluator disagreement explicitly. A smooth training curve cannot resolve an incoherent supervision contract.

Sources

Related field notes