Choosing Deterministic and Probabilistic Graders for Agent Traces
By DX Research Group · · Trace evaluation
Separate executable constraints from semantic interpretation and measure disagreement before using an LLM grader.
Some trading-agent errors have exact answers. An order either exceeds a stated maximum or it does not. Other questions require interpreting strategy language or deciding whether an explanation addresses the evidence. We would give those questions different graders, with explicit ownership for each outcome.
Let the executable rule keep its authority
A deterministic grader should read typed values and a versioned policy. It can check allowlists, size limits, required fields and terminal status. A semantic grader can evaluate whether a rationale faithfully summarizes supplied context or whether an instruction has an ambiguous interpretation. The semantic score should remain separate from the executable policy result.
The operating-layer controls account describes validation of typed actions and diagnosis of reasoning-trace failures. Those are distinct surfaces. A valid payload can carry an unsupported explanation, while a clear explanation can accompany an invalid payload. Combining both into a single success label discards information that helps repair the agent.
An illustrative action has a maximum permitted notional of 1,000 and a proposed notional of 1,050. The deterministic result is a violation under that policy version. A language-model grader might judge the action cautious because the rationale describes low risk. That interpretation cannot override the typed comparison. Store both results and expose the disagreement.
Treat the judge as another measured component
For a proposed semantic evaluation, write a rubric with examples that define each label. Freeze the judge model, prompt and input scope. Save its outputs and any reported probabilities. If the judge sees venue outcomes while scoring explanation quality, hindsight can contaminate its assessment of what the agent knew.
Review a sample stratified by score and disagreement, rather than only the easy high-confidence examples. Include contradictory actions, invented rules and legitimately ambiguous mandates. Human review should identify what the rubric means in those cases and measure agreement on that reviewed population. A raw probability emitted by a judge becomes a calibrated probability only after an appropriate calibration assessment.
Changing the judge prompt can change the apparent agent score. Keep judge versions in the result record and rerun the same saved traces when making a comparison. If old outputs remain in a mixed table, label them by judge version rather than treating them as one continuous measurement.
Keep the repair decision local
The continuous record emphasizes concrete baselines and held-out validation in its methodology. We apply that discipline to the evaluator too: a semantic judge should add useful information beyond simpler checks before it influences an engineering decision.
DXAP places policy checks outside the model. That architecture gives deterministic restrictions a clear role while leaving semantic review available for diagnosis. A proposed scorecard would show executable violations, explanation labels and completion separately. Engineers can then repair the order path, renderer or rubric responsible for a failure, and users can understand which kind of reliability a reported score actually measures.