Attribute Tool-Chain Failures at the Broken Dependency

By DX Research Group · · Trace evaluation

A dependency audit separates the first corrupted result from later calls that merely propagate it.

A failed final order can originate several calls earlier. We would attribute a tool-chain defect to the earliest evidenced broken dependency, while recording every downstream consequence. Counting the final error alone directs the repair toward the place where the failure surfaced rather than the place where it began.

A unit error travels through three calls

Consider an illustrative chain: fetch a quote, calculate position size, then construct an order. The quote connector returns cents where the contract promises dollars. A value of 10,000 represents a price of 100 dollars. The sizing tool accepts 10,000 as dollars and divides a 1,000-dollar budget by it, proposing 0.1 units rather than ten units.

The order constructor faithfully uses 0.1 units. A reviewer sees undersizing at the final stage, but the earliest evidenced divergence is the quote unit. Calling this an order-construction failure misstates the dependency. If the connector returns the right units and sizing divides incorrectly, the primary label moves to sizing.

OpenTelemetry's tracing API supports parent-child spans and links. Our proposed audit adds input-output dependencies, because asynchronous work can feed several children and the causal path can differ from a simple execution tree.

Preserve consumed values, not just tool names

For each call, retain the upstream result identifier, consumed fields and their units. A successful transport response can contain the wrong semantic value. The audit should distinguish availability, structural validity and domain correctness before judging downstream reasoning.

The operating-layer controls companion motivates full-path inspection. The continuous record companion provides historical multi-tool context. This note proposes a dependency-localizing method; it supplies no new estimate of connector failure prevalence.

Build a paired offline fixture by replacing only the corrupted quote with a contract-correct quote. Hold the remaining chain inputs fixed and observe whether sizing and order construction recover. If the model still chooses 0.1 units, the connector error was real but insufficient to explain the final action. Preserve both findings rather than declare one cause from a plausible narrative.

Branching chains need more care. A portfolio result may inform sizing and an exposure validator independently. Two visible errors can share one upstream cause, while a correct validator may block the malformed action. Record the dependency graph and first divergence per branch. Count impacted decisions separately from defective calls.

For the proposed release report, include one reconstructed chain with raw values, expected units, corrected substitution and residual outcomes. Mark unavailable dependencies unknown. The method's useful result is a repair hypothesis tied to an observed value transformation. It lets us test whether fixing one connector actually restores the intended decision path, instead of celebrating a lower downstream error count whose origin remains unexplained.

Sources

Related field notes