Research
What LLM Trading Agents Actually Do in Production: A Six-Month, Population-Scale Record from Two Fleets
What LLM Trading Agents Actually Do in Production: A Six-Month, Population-Scale Record from Two Fleets Overview Research area: Autonomous LLM agents operating in live financial markets; agent evaluat
- arXiv
- 2609.05663
- Published
- 2026-09-04
- Authors
- T. J. Barton, Chris Constantakis, Patti Hauseman, Annie Mous, Alaska Hoffman, Brian Bergeron, Hunter Goodreau
AI summary
What LLM Trading Agents Actually Do in Production: A Six-Month, Population-Scale Record from Two FleetsOverview
Research area: Autonomous LLM agents operating in live financial markets; agent evaluation and measurement methodology; market microstructure and risk behavior of AI trading systems.
Technical level: Intermediate. The paper is written in accessible operational language, but its evidence relies on regression discontinuity, day-clustered inference, Mantel–Haenszel stratification, permutation nulls, and paired replay leagues.
Scope in one sentence: A six-month, population-scale measurement record of two production fleets of autonomous language-model trading agents — DX Terminal Pro (3,505 user-funded vaults trading real ETH in Base memecoin markets for 21 days) and the DXAP live alpha fleet (500–599 user-created agents trading Hyperliquid perpetuals) — showing that the operating layer around the model, not the model's strategy text, determines behavior, and that neither fleet shows a directional edge.
What This Paper Is About
By mid-2026, autonomous LLM agents routinely move real money in public markets, but the field evaluates them mostly through leaderboard P&L over hours or days, with no sustained account of what thousands of such agents actually do when users hand them sliders, strategy text, and capital. This paper supplies that record: it measures a continuous six-month span across two production systems built on one design lineage, spanning 7.5M single-model invocations, about 300K onchain actions, and a further 231,638 multi-tool turns producing 14,596 fills. Its goal is not to demonstrate a profitable trading strategy — the authors state plainly that neither population shows a directional edge — but to localize which parts of the surrounding system determine agent behavior and where the defects that produce losses live.
Key Contributions
-
A continuous population-scale production record spanning two systems on one design lineage. The paper measures DX Terminal Pro (3,505 user-funded vaults, 7.5M invocations, ~300K onchain actions, 21 days, February to March 2026) and the DXAP live alpha fleet (500–599 agents all-history, 91–117 concurrently active, 231,638 finalized turns, 14,596 fills, June to August 2026), pooling at the metric level rather than the row level.
-
The identification of the operating layer as the dominant determinant of agent behavior. Configuration sliders, rendered candidate lists, and order-path mechanics explain behavior more than anything written in strategy text, including a clean causal regression-discontinuity result at a leaderboard render boundary.
-
A quantified diagnosis of the defects producing losses, including volatility-blind sizing across 6,400 closed positions, liquidation concentration in one posture-slider cell, and a capture gap in which 43.2% of positions reached at least +300 bps of favorable excursion within 24h yet 49.3% of those closed negative.
-
A 17-rule methodology canon for agent-trading research, each rule bought with a retraction or failed claim in the authors' own program, plus a development program ordered by the record's effect-size map.
Main Findings
-
The operating layer, not the strategy text, sets behavior. In DXAP, the
riskToleranceslider maps to leverage at +0.425× per level (p = 3.8 × 10⁻²⁷⁹), agent fixed effects absorb 60% of variance, and leverage named in strategy text is tracked at Spearman ρ = +0.836. Nothing else measured (market state, recent P&L, prompt family) moves sizing comparably. -
A rendering choice causally routes selection. DXAP agents see a nine-symbol "movers" leaderboard. 46.5% [43.4, 49.7] of entries are in rendered symbols against an 8.9% random-availability baseline (5.2× over-selection). A regression discontinuity across the rank-3/rank-4 cut gives a selection ratio of 1.75× [1.49, 2.06] exactly on the boundary; the raw rank-3/rank-4 count ratio is 1.81× (290/160). The authors call this the one clean causal result in the record.
-
Sizing is volatility-blind. Across 6,400 closed positions split into volatility sextiles spanning a 5.7× spread, median leverage is 5.0× in every sextile; Spearman(volatility, leverage) = −0.001 (p = 0.92). Notional goes the wrong way: Spearman(volatility, notional) = +0.165 (p = 2.7 × 10⁻⁴⁰). Median realized return degrades from −10.6 bps in the calmest sextile to −98.2 bps in the wildest, and liquidation rate rises from 0.7% to 4.3%. Excluding all 205 liquidations, return still falls from −8.5 to −53.6 bps (ρ = −0.104, p = 1.8 × 10⁻¹⁶, n = 6,195).
-
Liquidation risk is a setup-time configuration choice, not a series of bad decisions. Of 205 liquidations among the 6,400 closed positions (3.2%, carrying 74% of gross loss), 128 (62%) sit in a single cell: momentum-posture agents at frequency-slider 5, roughly 11% of the book. The day-stratified Mantel–Haenszel odds ratio is 22.37 [12.59, 37.45]. The cell spans 70 distinct agents, and across 411 momentum-posture positions at frequency below 5 there were zero liquidations.
-
Forced numeric restatement fails. Requiring the model to compute and state its liquidation distance before entry did not restrain it: agents state liquidation distance in 45.0% of entry turns, size identically anyway, and staters liquidate more often (5.8% vs. 1.2% for non-staters), because stating marks aggressive intent rather than restraining it.
-
A large capture gap. 43.2% of closed positions (2,765 of 6,400) reached at least +300 bps of maximum favorable excursion within 24h, yet 49.3% of those closed with a negative trade return; only 13.3% kept half the excursion. Book-wide median favorable excursion is +247.9 bps against a median realized −33.3 bps; median capture where upside existed is 2.0%. The market state predicts what the position will do (ROC 0.751 for upside, 0.783 for liquidation) but not what the agent will keep (ROC 0.541, essentially chance).
-
A mechanical bracket beats discretionary exits. Attaching a fixed 2%/4% stop/target bracket at entry earns +39.0 bps per position [+21.3, +56.5] paired and day-clustered; excluding all liquidations it still earns +16.6 [2.1, 30.8]. Roughly 60% of the gain is blow-up prevention (+717 bps across the 205 liquidations). 11 of 16 exit policies clear zero; none is profitable outright (best −9.5 [−25.1, +8.2]).
-
Neither fleet shows a directional edge. The DXAP fleet runs a 41% roundtrip win rate (n = 353) against Hyperliquid retail's 50% (n = 16,009) over the trailing 4-day window ending Aug 15, at nearly identical median hold times (2.05h vs. 2.18h). Only 15% of week-active agents are net-positive versus 53% of retail accounts. Cumulative realized P&L is −$217K at a common 5.5 bps fee rate, against −$148K at zero fee. No user was net-positive at the Jul 12 census.
-
The agents' one relative edge is horizon, not skill. Positions held 24h+ earn +0.23% median account return while sub-1h positions earn −0.15%.
-
DX Terminal Pro showed bursty, social, loss-making behavior. 1,544 of 3,454 active vaults bought the same token (FEET) within one hour on Mar 1 (peak 441 buys/min at 19:04 UTC); of the 908 launch-day buyers, the 82 still holding at measurement were 89% underwater (mean P&L −20.1%). The system recorded 3,878 sell cascades and 92.9% of trades occurred in two-sided 5-minute windows. Median true return was 0.492×; 16.2% of vaults finished profitable; 92% of deposited ETH was withdrawn.
-
Structured control surfaces outperform conversation. DX Terminal Pro owners who wrote concrete, numeric instructions were profitable 4.2× as often as the median, and the 87 UI-only owners who never used chat were the highest-profit cohort (41% in profit). A cross-system contrast sharpens this: in DX Terminal Pro, sliders override strategy text; in DXAP, strategy text routinely overrode the frequency slider. Which surface wins is an operating-layer decision.
-
Frontier models are indistinguishable in decision quality at this horizon. A paired replay league of 416 captured production scenarios across 7 days at temperature 0.6 finds three frontier models spanning 263.27–264.37 bps of winsorized regret, with every interval overlapping every other and Holm-adjusted tests against the leader all non-significant (minimum p = 0.46). Choice stability does differ: in the forced-action arm (280 cells), claude-fable-5 changed its chosen (symbol, side) action in 35% of cells against ~90–95% for the qwen3.7 pair, but this variation carried no measurable economic cost (selection quality spans rank percentile 0.486–0.508). Economics do separate them: qwen3.7-plus completed the 416-cell league for $8.25 against $203.61 for claude-fable-5, a roughly 25-fold cost difference.
-
The major nulls. A 77K-candidate signal ledger sorts at chance (leave-one-out win 0.50); white-box probes over 689 traces add +0.001–0.003, at or below the permutation null; 770K tokens of context yields −0.46 points [−2.93, +1.75]; a world-context arm went 0/84. Across 8 strategy postures, 6 prompt templates, and 9 cohorts, no group's lower day-clustered confidence bound exceeds zero. Memory-write frequency correlates negatively with P&L (ρ = −0.200, p = 0.002), and one agent ran 31 of 32 positions under strategy text that had been deleted.
-
The preregistered H2 risk-representation probe is narrowly positive and heavily qualified. A small encoder reading the rendered turn card beats concrete features by +0.0150 PR-AUC (90% day-clustered CI [+0.0013, +0.0312]), but the anchor ladder shows the increment appears only after the order line is read; pre-order anchors add nothing. At a 10% flagging budget the representation catches 71/205 liquidations vs. 57/205 for concrete features (precision 11.1% vs. 8.9%).
-
An early internal training signal. DX-SFT-0.3, a supervised fine-tune on 41,203 harness turns, improves on the lab's internal DXTradeBench decision evaluation from 119 (baseline) to 96 (lower is better; internal lab report, early, not externally validated).
Methodology in Plain English
The authors did not run a controlled laboratory experiment. They instrumented two live products that real users were already using, and then mined the resulting telemetry for regularities.
Two systems, one lineage. Both systems expose the same five-slider configuration surface (trade activity/frequency, risk tolerance, trade size, holding style, diversification; each 1–5) plus free-text strategy with priority and expiry. Both run a one-action-per-turn loop in which the model receives a compiled context and must emit typed tool calls with rationale. Both compile prompts from Go templates whose version lineage runs continuously from DX Terminal Pro's v2.x series into the DXAP template families.
Data stores and join discipline. DX Terminal Pro telemetry lives in Athena, reconciled against onchain vault state. DXAP telemetry lives in a ClickHouse analytics store with paper fills and realized P&L, per-turn inference logs, valuation snapshots at roughly 2.5-minute cadence, and trigger/memory/config-revision tables. Historical attribution joins on turn id + attempt id + the template and config revision resolved at that turn, never the current roster template.
Pooling. Because the two systems share a design lineage rather than a schema, the authors pool at the metric level (slider-to-behavior mappings, trade rate, hold time, win rate, fee burden) rather than at the row level.
Fee restatement. Because the builder fee moved across eras (4.5 → 14.5 on Jun 22 22:00 UTC → 5.5 bps/side), every cross-era P&L figure is restated at a common 5.5 bps rate.
Evidence classes. Every headline claim is marked FIRM (registered analysis on the full window, day-clustered uncertainty intervals, and at least one ablation or permutation check) or PROVISIONAL (positive but narrow, qualified, or sensitive to a design choice the authors could not fully close). Three earlier results in the program's history were retracted — a HIP-3 lag edge that turned out to be a timestamp artifact, and two "only net-positive posture/edition" claims killed by calendar-footprint checks — and appear only as retractions.
Causal identification. The one clean causal result is a regression discontinuity across the rank-3/rank-4 render boundary in the movers leaderboard, where symbols just below the cut are statistically identical in market state to those just above but are picked far less because they are not shown.
Benchmarking. A paired replay league scores frontier models as winsorized regret against an achievable envelope over 416 captured production scenarios, with day-to-8h-block hierarchical bootstrap intervals.
Stated limitations of the measurement environment. Most DXAP fills are paper: fills at live mark with zero slippage, zero funding paid across all 14,596 fills, and a 0.4% maintenance-margin placeholder that liquidates roughly 2× later than real Hyperliquid margining. The 5,035 real-money fills are flagged wherever they overlap an analysis rather than pooled silently or excluded. Liquidation-related results should therefore be read as earlier and larger under real margining.
Why This Matters
Impact on research. The paper reframes how agentic trading systems should be evaluated. The dominant practice — leaderboard P&L over short horizons — is shown to be uninformative about the mechanics that actually produce losses. The authors' own evidence is that decision quality across frontier models is statistically indistinguishable at the tested horizon (regret spread 263.27–264.37 bps, all intervals overlapping, minimum Holm-adjusted p = 0.46), while harness-level mechanics produce large,
Authors’ abstract
We present a continuous, population-scale measurement record of autonomous language-model trading agents operating in production across two systems with one design lineage: DX Terminal Pro (3,505 user-funded vaults trading real ETH in Base memecoin markets for 21 days, February to March 2026) and the DXAP live alpha fleet (500 to 599 user-created agents all-history, 91 to 117 concurrently active, trading Hyperliquid perpetuals, June to August 2026). The record spans roughly six months, 7.5M single-model invocations with about 300K onchain actions, and a further 231,638 multi-tool turns producing 14,596 fills. Four findings carry the paper. First, the operating layer determines behavior more than anything written in strategy text: a risk slider explains leverage (+0.425 per level), agent fixed effects absorb 60% of variance, and a leaderboard render boundary causally routes selection (regression discontinuity 1.75x at the top-3 cut). Second, sizing is volatility-blind: median leverage is 5.0x in every volatility sextile, and one posture-slider cell (11% of the book) holds 62% of liquidations. Third, agents capture almost none of the upside they reach: 43.2% of positions saw at least +300 bps of favorable excursion within 24h, yet 49.3% of those closed with a negative trade return; a mechanical bracket recovers +39.0 bps per position. Fourth, neither fleet shows a directional edge. The DXAP fleet is not profitable and trails a matched Hyperliquid retail benchmark (41% vs. 50% roundtrip win rate). A paired-replay league of frontier models on 416 captured production scenarios finds decision quality statistically indistinguishable at this horizon, while choice stability differs sharply across model families. Every headline survives day-clustered inference, permutation nulls, and a common-fee restatement; the paper closes with a 17-rule methodology canon bought with our own retractions.