Separating Prompt Order from Context Length in Trading-Agent Ablations

By DX Research Group · · Learning theories

A four-arm trading-agent experiment crosses assigned context length with instruction placement to distinguish salience changes from action improvements.

If a trading agent improves after we shorten its input and move a risk instruction forward, we have changed two things. The experiment cannot tell us whether the improvement came from length, position, or their interaction. We would cross those factors explicitly before assigning the result to either one.

The proposed study uses the same decision-relevant facts in four conditions. It tests how the agent handles an instruction under different surrounding lengths and positions. This is a prompt ablation, rather than a comparison between a compact state snapshot and an entire trading history.

Four Arms With One Question

Choose two assigned input budgets within the selected model's verified native window. Modern 256k-plus capacities can support large experiments, but an assigned token budget is a study setting, not a statement about the model's native capacity. Verify native support and serving limits separately before preparing packets.

Cross the shorter and longer assigned lengths with early and late instruction placement. The four arms are short/early, short/late, long/early, and long/late. Each case retains identical market facts, portfolio values, mandate constraints, and tool outputs. Move the target instruction without changing its words.

For the longer condition, add a declared set of surrounding material that supplies no additional answer to the target question. Preserve that material across the early and late arms. Document its content, since realistic commentary and repetitive padding can produce different effects. The test therefore estimates a length effect under the chosen surrounding material, with that scope carried into the result.

Measure Salience Before Inferring Decisions

Our published controls study supplies a narrow precedent. Moving an unchanged fee sentence from paragraph eight to paragraph one raised fee citation in sampled reasoning traces from 3% to 74%, with the model, sentence wording, and market data fixed. Exact per-arm counts are absent from the published account. The observation concerns trace citation, with economic gains left unresolved.

Use that distinction in the new experiment. Record whether the agent mentions the target instruction, interprets it correctly, and lets it affect the proposed action. A citation score measures visible salience. A compliant action measures behavior under the mandate. Neither alone establishes better returns.

An illustrative case has $500 of existing exposure, a $700 cap, and a proposed $300 addition. The correct remaining capacity is $200. Score recognition of the cap separately from whether the agent proposes an admissible size or observes. All four arms must expose those same numbers.

Pair Cases Across the Matrix

Evaluate every saved case in all four conditions with model version, sampling settings, and policy checks fixed. Preserve paired outcomes so an average cannot conceal opposite effects on different cases. Report the early-versus-late difference within each length and the longer-versus-shorter difference within each position.

If early placement helps only in the longer condition, the interaction is the finding. If citation rises while actions stay unchanged, the intervention improves the trace diagnostic alone. Our harness-transfer method provides the component attribution discipline; the trace feedback framework explains how to preserve these outcomes as regression cases.

DXAP describes a data index, model proposals, policy checks, and recorded turns. That architecture makes prompt experiments relevant as the harness evolves. This four-arm study remains proposed, and any release decision would need its paired results and downstream checks before we claim an improvement.

Sources

Related field notes