Choose a model update from the problem users actually report
By DX Research Group · · Data and learning flywheels
A proposed problem taxonomy routes curated user contributions to a specific training objective and held-out test.
The most valuable next model update depends on what is failing for users. We propose classifying reviewed user problems before selecting a training objective. The classification should identify the behavior a model can improve, the data that would teach it and the test that could show the improvement.
The controls paper companion describes tracing failures to stages in the mandate-to-settlement path and testing harness revisions against replayed scenarios. That historical method motivates a user-contribution pipeline with the same diagnostic discipline. A complaint is evidence of a problem; its frequency alone does not establish that changing model weights will repair it.
Take a fictional week of feedback. Owners report unexplained waiting decisions, misunderstood position limits and missing order-status updates. Review finds that some waiting explanations bury the useful condition, some limit questions confuse notional with margin, and some status updates are absent because a tool response failed. These cases need different destinations.
The first group can produce presentation preferences. The second can produce supervised examples grounded in DXAP’s configuration reference, which defines notional as exposure rather than margin or maximum loss. The third initially needs a tool or delivery investigation. Teaching the model a polished status message cannot supply an absent execution fact.
Select the lesson before the optimizer
An illustrative intake has forty reviewed cases: eighteen presentation problems, twelve concept misunderstandings and ten missing-tool incidents. If eighteen presentation cases came from two highly active owners, their count alone should not decide a shared update. Retain affected-owner counts, severity and repeat-incident identity alongside the case total.
For a concept objective, a curated row contains the owner question, supported configuration state and expected explanation. For a presentation objective, it contains fact-matched alternatives and an owner judgment. For a tool failure, it contains an incident fixture and recovery expectation. That third record may later teach a truthful incomplete-state response, once the infrastructure defect is understood.
We would choose an objective using the expected reduction in a defined user problem, label reliability and whether the proposed model change addresses the diagnosed cause. The choice should state why a prompt change, documentation repair or deterministic check would be insufficient or less suitable. A taxonomy organizes options; it should leave room for a combined repair where the evidence supports it.
Give each objective an exit test
Suppose the selected objective is explaining notional correctly. Hold out owner questions and numerical examples that differ from training. Ask the candidate and baseline to explain the same saved states. Score conceptual correctness, useful calculation and invented settings claims. Include an adjacent case where the owner asks about maximum loss, requiring an explanation of why the notional setting does not answer that question.
The chat guide makes approval boundaries relevant to that test: an explanation or proposal must preserve the separate owner confirmation step. A model that answers the concept correctly while claiming it changed the limit has introduced another user problem.
The proposed release decision would attach the selected objective, curated example types and held-out results to one update record. Unresolved classes stay visible for later work. This turns participation into a route from reported friction to a measurable training objective, while preserving the option to fix the harness when the harness owns the failure.