What happens when an owner withdraws a contribution?

By DX Research Group · · Data and learning flywheels

A proposed dependency graph makes withdrawal concrete across source records and derived datasets.

Withdrawing a contribution becomes tractable only when the system knows where that contribution went. Removing the original upload leaves copied fixtures, dataset exports and evaluation summaries untouched unless they have explicit dependency records. We propose a withdrawal workflow built around artifact lineage and consumer receipts. The public sources inspected here establish no current DXAP service implementing this workflow.

The controls paper companion links recorded stages so researchers can attribute an action to its inputs. A contribution pipeline could extend that discipline to source report, transformed fixture and dataset release. Each derived object would record the source versions it used and the permission receipt checked at creation.

Follow one example through its descendants

Consider an illustrative contribution R. A curator turns it into synthetic fixture F. Dataset D includes F, and evaluation E uses D. An owner withdraws R before a planned training run. The system first prevents new exports from R and F, then creates a revised dataset D2 excluding F. The planned run must resolve D2 explicitly or stop; keeping a cached copy of D under the same filename defeats withdrawal.

The system should preserve a minimal restricted tombstone showing that an artifact was withdrawn and which descendants require action. That record need not retain the sensitive content. Deleting the dependency identity immediately would make it harder to verify exclusion from later releases. Access to the tombstone should itself have a defined operational purpose and retention period.

For completed evaluation E, the response depends on what it contains. A public aggregate with no recoverable example content may need an amended provenance note or recomputation under the declared policy. An exported dataset containing F needs removal from controlled distribution and a notice to recorded recipients. Receipts establish who received it; they cannot guarantee retrieval of copies outside the operator's control.

Dataset removal and trained-model effects differ

If a completed training run consumed D, exclusion from future datasets is a concrete action. Removing the contribution's effect from existing model weights is a separate technical problem. A proposed operator should report the affected checkpoint and available remedy, such as retraining from an earlier uncontaminated checkpoint, instead of calling file deletion model unlearning. Any unlearning method would need its own verification and a stated uncertainty about residual influence.

We would test withdrawal before export, after internal export and after a completed evaluation. For each scenario, the review checks blocked future use, identified descendants and updated release manifests. A deliberately unregistered copy should produce an explicit lineage failure, exposing the coverage limit of the system.

This design gives contributors a concrete view of what withdrawing can change. It also forces researchers to distinguish planned, distributed and already-consumed artifacts. Its first measurable result would be complete descendant reconciliation within the controlled pipeline. Whether a contribution once improved instruction compliance, forecasts or decision economics remains a separate empirical question, and withdrawal should preserve that distinction in the historical record.

Sources

Related field notes