research

Can ChatGPT Trade? What the Evidence Shows

How a ChatGPT trading bot connects to execution tools, and what the published record shows about performance and account controls.

By DXRGPublished Updated Follow research via RSS
Can ChatGPT Trade? What the Evidence Shows

ChatGPT can help you research markets, review information you provide, and make a trading idea easier to inspect. Placing trades is a separate capability: it depends on the app, connected tools and permissions you have actually enabled. Robinhood, for example, documents a connection for ChatGPT and other compatible AI platforms through its Agentic Trading product.1 A connection tells you what the system can access and do; its trading results need their own evidence. This page explains useful tasks, what control you retain, and what the published record can establish.

Evidence boundary: This page summarizes a third-party live competition, a third-party benchmark preprint, and DXRG controlled tests measured on a different model family. Every figure is bounded to the setting that produced it, and no result here establishes profitability, universal safety, or future performance for any model.

The Short Answer

Three claims get collapsed into one whenever someone asks whether ChatGPT can trade, and they have different answers.

The first claim is connectivity. Can a ChatGPT-connected system reach a venue and place an order? A compatible, authorized integration can provide that path. Robinhood's Model Context Protocol (MCP) connection exposes account information and trading tools to supported AI platforms. Its ring-fenced Agentic account limits where the agent can trade; it does not limit all the information the connection can read. Approval requirements also depend on the setup, as explained below.1 For a ChatGPT trading bot, verify the tools and authority actually available in your session instead of assuming the product name settles either question.

The second claim is competence. Does the raw model, handed market data and left to decide, show a demonstrated edge? The cited studies do not establish a repeatable raw-model edge in their measured settings. Their scope and results are below.

The third claim is control. When a trade happens, what governed the exact action that reached the market? That is the question our work answers, and the answer is the harness: the system around the model that compiles the mandate, assembles state, checks policy, executes, and records the trace. Two setups running the same chatbot can produce opposite outcomes because their harnesses differ.

What can you hand to an AI trading assistant?

Start with the work you want done. OpenAI's current documentation describes ChatGPT Work searching the web, comparing sources, reading files and analyzing data, with availability depending on the plan and workspace.2 Those capabilities can support a research task before you grant a service access to a trading account.

Your taskA useful outputWhat you still check
Understand a market moveA brief connecting dated announcements to the question you askedOpen the original sources; distinguish reported facts from the assistant's explanation
Review your portfolioAn explanation tied to the holdings and account snapshot it can actually seeConfirm which account and data it used; decide whether any change fits your circumstances
Turn an idea into rulesA specification with missing assumptions made explicitReview the rules and test them before relying on their results
Run an ongoing trading strategyDecisions and actions through an authorized trading systemCheck the allowed account, trading limits, approval requirements and how to stop it

For the first task, a research request could be:

Compare the company's latest earnings release with the previous quarter. Link the original releases and their publication dates. Separate reported figures, management forecasts and your interpretation. List missing information. Produce a research brief without placing orders.

This is an example request, not a tested strategy. Review the cited documents before acting on the answer. For the third task, our strategy-specification example shows how to turn an incomplete idea into questions and test cases; there is no need to start with an automated trade.

An account-aware assistant can reduce the work of bringing your holdings into the discussion. Robinhood's August 25, 2026 description of Cortex gives portfolio questions and market research grounded in account data as examples.3 Cortex is its in-app assistant; Agentic Trading connects a third-party AI platform to a dedicated trading account. Check the particular product's capabilities rather than treating these as interchangeable names for an autonomous trader.

Will it ask before placing each trade?

Do not assume that every order requires another click. Robinhood's current help page says an agent can place trades without your confirmation if you have asked it to act without requesting approval.4 Your AI platform may apply additional restrictions. Check both the platform and the broker settings for the workflow you intend to use.

There is a separate privacy distinction. Robinhood documents read access to all your Robinhood accounts and their positions, balances and transaction history, while limiting trading to the Agentic account.1 A separate trading account therefore limits where orders can be placed, without necessarily limiting which account information is shared. Review data access and trading authority as two different choices.

For an ongoing agent, ask how it runs after you leave the chat and how you can pause or revoke it. A written instruction is useful only to the extent the running system follows or enforces it. We use the evidence below to evaluate that system, rather than treating a fluent answer or a successful connection as a return result.

What the Public Record Shows

The cleanest public measurement so far is Alpha Arena Season 1, run by the research lab nof1.ai.5 Six frontier language models each received $10,000 in real capital and traded crypto perpetuals on Hyperliquid from late October into November 2025, a window of roughly two weeks. Every model saw the same market data format and the same action space. The results, sorted:

ModelSeason 1 result
Qwen3-Max+22.3%
DeepSeek+4.9%
Claude−30.8%
Grok−45.3%
Gemini−56.7%
GPT-5−62.7%
Alpha Arena Season 1 results sorted by return: GPT-5 finished last at −62.7%Alpha Arena Season 1 results sorted by return: GPT-5 finished last at −62.7%

GPT-5 finished last in the field, down 62.7% in about two weeks. Three of the six models lost a third or more of their capital, and the two positive results came from the same rails that produced those losses. The spread between first and last place was 85 percentage points across identical wiring.

That result deserves a careful reading. Alpha Arena is an anecdote-class measurement in the strict sense: one season, one venue, one prompt setup, two weeks. It records what happened when six models traded real money under one fixed configuration. It says little about any model's permanent ranking, and a second season could reorder the table. What it does establish is narrower and still useful: identical wiring plus different models produced a wide outcome spread, and most of the field lost money quickly. The arena made no attempt to mask future information, vary prompts, or separate luck from selection, and it ran once. None of that makes the result meaningless: real capital moved on a real venue, and the full field traded the same two weeks of market. It makes the result one observation per model, closer to a race photograph than to a timed qualifying series. Anyone selling a ChatGPT trading bot should be asked which measurement their claimed edge survives.

The second public source is KTD-Fin, a memory-controlled benchmark for LLM trading agents on stock markets (arXiv:2605.28359v1).6 Its cross-model design holds prompts, tool interfaces, masking, and execution rules constant. The benchmark anonymizes tickers and calendar identifiers across prompts and tools, then separates each model's result into common, style, and stock-selection components. The paper reports that 9 of 10 frontier LLMs showed negative stock-selection alpha, meaning their picks did worse than the style exposures they took on. Our benchmark audit registry now records a source-disclosure review with all eight units marked partial. We still attribute the result to the authors because our review did not rerun the benchmark, validate its data or results, or certify the masking protocol.

These two sources measure different things, and they complement each other. Alpha Arena is live evidence: real capital, a real venue, a short window, and no controlled masking. KTD-Fin is a controlled benchmark: historical markets, identifier-and-calendar masking, and component attribution. One records what happened in a public two-week competition; the other measures whether each model's picks beat the exposures those picks took on. Both point in the same direction: across the measured public settings, no frontier model, GPT-5 included, has shown a repeatable raw-model edge.

What the Harness Record Adds

The strongest evidence we hold on what governs outcomes comes from our own controlled record, and we want to be precise about its scope: those numbers were measured on the Claude model family. We report them here as harness evidence, a measurement of what the surrounding system changes, and we claim no transfer of the specific figures to GPT models. What transfers is the mechanism. We include these numbers on a ChatGPT page because the harness effect is the part of the record we can measure directly, and because the gap it exposes, model text on one side and governed action on the other, is the gap every chatbot-to-broker connection has to cross.

In a separate internal EVM transaction-construction evaluation, aligned successful construction was 87% for Claude 4, 96% for Claude 4.6, and 99.9% for Claude 4.6 with a Terminal Pro-style harness.7 The model generation moved the number by nine points. The harness around the stronger model moved it by another four, to within a tenth of a percent of perfect construction.

A harness is everything around the model that turns text into a governed action. In our stack that means compiling the owner's mandate into an explicit contract, assembling current market and portfolio state, and invoking the model. It then means parsing the answer into a typed action, checking that exact action against deterministic policy, submitting once, and reconciling the outcome into the next state. Each stage is recorded in one trace. When someone says ChatGPT traded, the useful follow-up is which of those stages the output passed through before money moved.

Two controlled pre-launch tests from the DX Terminal Pro record show the mechanism at finer grain. Moving an unchanged fee sentence from paragraph eight to paragraph one raised fee citation in reasoning traces from 3% to 74% while the model, the wording, and the market data stayed fixed.7 A compound intervention that demoted prior reasoning from precedent to context reduced fabricated sell rules from 57% to 3% in the affected test population.7 Same model, same facts, different compiled context, different behavior.

These figures measure construction quality and behavioral control on one model family. They sit apart from open-market P&L, and we report them with the same framing we used in our FSB consultation response and in the paper companion. For a ChatGPT-connected setup, the open question is empirical: run the same harness gradient on the model you intend to connect and measure where it lands. Our guardrails article describes that system in engineering form, and the guardrail matrix dataset (CSV) maps each control to its evidence requirement and failure signal.

What to Check Before Connecting Any Chatbot to Money

Whether the model is GPT-5, Claude, or anything else, the evidence above reduces to four checks before capital moves:

  1. Custody. Who holds the keys, and what can the model's credential actually do? A scoped API key that can trade but can never withdraw is a different risk from an account-level session. A ring-fenced brokerage subaccount is a different risk from the account that holds your salary.
  2. Proposal versus decision. Does generated text reach the venue directly, or does deterministic software outside the model check the exact action against policy first?
  3. Evidence class. Is the claimed result simulation, replay, paper, shadow, or live, and does the claim match the class? Our evidence ladder defines the rungs.
  4. Disclosure. Does the setup answer the twelve questions in our benchmark card, including costs, baselines, repeated runs, and the trace from instruction to outcome?

The broker case adds one more question: what does the integration let the model do by default, and what still asks for your confirmation? Read a trading permission screen the way you would read a margin agreement.

Those checks help you understand the system you are using. They do not establish whether its strategy will earn a return.

The Bottom Line

Can ChatGPT trade? The rails exist: since May 27, 2026 a major US broker connects ChatGPT to a ring-fenced brokerage account through MCP, and exchange APIs make the same wiring possible outside a broker. Does the raw model bring an edge? The cited studies do not establish that edge: in the one real-capital public season, GPT-5 finished last at −62.7%, and in the one memory-controlled stock benchmark, 9 of 10 frontier LLMs showed negative stock-selection alpha. What changes outcomes is the system around the model: in our controlled record, measured on the Claude family, construction success moved from 87% to 99.9% across a model generation plus a harness, and the position of a fee sentence in the rendered context moved citation rates from 3% to 74% while everything else stayed fixed. This article is the second in an evidence series; the first covers Claude, and the harness figures above come from the record behind our paper companion. Evaluate the whole path, and treat any unstated evidence class as unknown until the supporting record is disclosed.

Version 1.3. Published August 17, 2026; updated September 9, 2026. This revision adds practical assistant tasks and corrects the account-access and approval explanation against current provider documentation. Historical research figures were not rerun. Corrections: poof@dxrg.ai. This article is research and educational material, not financial advice.

Footnotes

  1. Robinhood, Agentic Trading overview, accessed September 9, 2026. See “What your agent can access” and “Connect your AI agent.” (Robinhood help) 2 3

  2. OpenAI, Use ChatGPT, “What ChatGPT Work can do,” accessed September 9, 2026. (Official documentation)

  3. Robinhood, Robinhood Cortex: The Loop Stays Dumb So the Model Can Be Smart, August 25, 2026. Product description, not independently measured investing performance. (Robinhood engineering article)

  4. Robinhood, Trading with your agent, accessed September 9, 2026. See “Safety first.” (Robinhood help)

  5. nof1.ai, Alpha Arena Season 1, six frontier models trading $10,000 each in real capital on Hyperliquid, late October into November 2025. Results as published by the organizer. (nof1.ai)

  6. From Knowing to Doing: A Memory-Controlled Benchmark for LLM Trading Agents on Stock Markets (KTD-Fin), arXiv:2605.28359v1, 2026. DXRG completed a source-disclosure review on August 18, 2026; all eight units were partial, and DXRG did not rerun or validate the result. (arXiv abstract)

  7. Barton, T.J. et al., Operating-Layer Controls for Onchain Language-Model Agents Under Real Capital, arXiv:2604.26091, 28 April 2026. (arXiv abstract) 2 3