Detect Prompt Truncation Before Scoring Agent Decisions
By DX Research Group · · Trace evaluation
A section receipt and boundary-marker fixture identify missing instructions before an agent trace receives a behavioral grade.
Prompt truncation should be diagnosed as an input-integrity problem before we judge the agent's decision. A stored full prompt can differ from the serialized request and from the material admitted by the serving interface. We would retain evidence at each boundary and flag cases whose operative instructions may have disappeared.
An exception falls beyond the boundary
In an illustrative fixture, the strategy permits adding exposure, but a final paragraph temporarily prohibits opening a position in one instrument. The local prompt file contains both sentences. A client-side budget reducer removes the final paragraph while preserving the strategy summary. The agent opens that instrument. Against the complete file, the action violates the mandate; against the delivered request, the exception was absent.
The first task is therefore to reconstruct delivered sections. Give each section a stable identifier, byte length and local digest. After serialization, record which sections survived and any intentional compaction. Preserve the rendered request in the controlled evaluation environment, subject to its privacy policy. A digest alone shows equality between saved objects; it reveals little about the meaning of omitted text.
JSON's interchange specification helps define the serialization boundary. Our proposal adds semantic section receipts because syntactically valid transport can still carry an incomplete mandate. Token estimates should name the tokenizer and include tool schemas or other interface overhead where applicable.
Probe both ends and the middle
We would construct a length sweep with a short critical restriction at the beginning, middle and end of otherwise comparable context. The fixture asks the agent to return a receipt for supplied section identifiers before producing its typed proposal. Missing identifiers are a diagnostic signal, while exact request capture remains stronger evidence of what the client sent. A model can also fail to report an identifier that it received.
For a numerical illustration, a request has six required sections and two optional histories. The reducer removes both histories, leaving six of six required sections. Another reducer preserves seven sections but removes one required restriction: the count looks larger while mandate coverage falls to five of six. Our audit measures required-section coverage separately from total length.
The operating-layer controls companion ties behavior to compiled instructions. The continuous record companion makes the historical rendered context relevant to interpretation. Neither publication establishes the outcome of this proposed truncation fixture.
A flagged case should remain in the interface-failure report and leave the clean decision-quality slice until its input status is resolved. We would inspect client reduction, gateway limits and documented provider admission behavior in that order. The final artifact is a section-by-section receipt, with an explicit unknown where server-side visibility ends. That keeps a missing instruction from becoming a confident story about the model's risk preferences.