Quant engineering automation advances through verified changes
By DX Research Group · · Quant work and open markets
A task-based timeline for software, data and ML engineering explains why implementation speed and production ownership diverge.
Quant engineering should become one of the earliest areas where agents complete substantial bundles of work. The reason is practical: code can be executed, data transformations can be checked and performance can be measured. We expect the strongest adoption where teams can make a proposed change reviewable before it touches a live trading system.
The roles contain different verification problems. Jane Street's software engineer description covers systems across the firm, including network monitoring and risk models. Its data engineer description emphasizes understanding unfamiliar datasets and producing reliable research inputs. The machine-learning page separates model research, implementation and performance engineering. Our timeline uses these as task references, with dates that are explicitly DXRG scenario assumptions.
A migration that almost succeeds
Take an illustrative feed migration. An agent rewrites the parser, updates schemas and passes ordinary tests. During historical replay, one exchange message uses a timestamp with a different meaning. The transformation produces a clean table with information apparently available before it really arrived. The code is functioning, while the research input has become misleading.
This example explains why raw implementation speed tells us little about the whole job. Software correctness, data meaning and operational ownership require different checks. A useful engineering agent must preserve those distinctions as it moves between tasks.
For 2026-2028, our central scenario is agent ownership of scoped implementation packages: adapters, internal tools, documented migrations and regression repairs. Humans provide the intended behavior, permitted access and review criteria. The agent produces the change, relevant verification and a clear account of remaining uncertainty. In quant development, that can reduce the effort of turning a research specification into tested code.
The main adoption requirements are executable environments, representative fixtures and an independent review of the acceptance criteria. High confidence applies to changes whose failures appear in tests or replay. Confidence drops for performance changes judged on a quiet development machine or data changes whose semantics are understood by only one person.
Longer work needs better environments
The METR TH1.1 update expanded its software-task suite to 228 tasks and increased coverage of tasks lasting at least eight human hours. It also reports wide uncertainty and estimated rather than measured human times for many long tasks. The finding supports taking multi-step engineering autonomy seriously while keeping benchmark precision and workplace transfer explicit.
For 2028-2031, our scenario is agents maintaining bounded subsystems over repeated changes. A data agent could ingest an approved new source, compare it against an existing source, investigate discrepancies and maintain the resulting pipeline. An ML engineering agent could reproduce a model, investigate training bottlenecks and propose optimizations with equivalent-output checks. A software agent could carry a migration across multiple repositories under a defined dependency plan.
The bottleneck becomes verification coverage. If an agent can modify the application and the test that judges it, a passing result may simply reflect an easier test. We would keep critical acceptance fixtures outside the agent's modification scope and add unseen incident replays. Evaluate maintained behavior over time, including failures discovered after the apparent completion date.
A useful comparison holds the task specification, repository snapshot and available compute fixed. Count accepted changes, review hours and post-release defects. Track the severity of escaped failures instead of letting many harmless successful edits overwhelm one serious accounting error. Those are proposed measurements, with no completed DXRG result implied.
Production responsibility after 2031
From 2031 onward, continued capability gains could support highly autonomous engineering departments for instrumented systems. The full role still includes deciding which systems to build, coordinating with other owners and responding when the firm's assumptions change. Data engineering adds access rights and vendor semantics; ML performance work adds hardware behavior and numerics. Each has a different path to trustworthy delegation.
Incumbents have reasons to move quickly. An established codebase, representative internal data and detailed operational history can make an agent far more useful than the same model in an empty repository. Integration costs can slow deployment, while those assets can make successful deployment more valuable. Open tooling lowers the entry cost for challengers without erasing the value of a mature production environment.
We expect the profession to shift toward defining invariants, designing verification and choosing changes with economic value. The bold projection is that extensive hands-on implementation will become optional for many scoped quant engineering tasks this decade. Full production ownership will advance according to the quality of the surrounding evidence, with separate acceptance standards for code, data and model behavior.