Prediction markets allow users to trade on outcomes of real-world events, but are prone to fragmentation with overlapping questions, implicit equivalences, and hidden contradictions across markets. We present an agentic AI (AAI) pipeline that autonomously recovers cross-market structure from contract text before prices enter the analysis. The workflow first clusters markets into coherent topical groups using natural-language understanding over contract text and metadata, and then identifies contracts within each cluster, but from different event markets, that exhibit strong dependence or leader--follower relationships. We evaluate this system, along with a natural language inference (NLI) benchmark, on a large prediction market dataset from early 2026. Using resolved outcomes to evaluate identified relations, we find that AAI-identified relations are 62.8% consistent with exchange-recorded settlements, whereas the NLI benchmark only achieves 40.6% accuracy. Within clusters, the AAI output is sparse and also remarkably compatible as a signed graph with a frustration rate of 0.324%. As an application, we show how discovered relations inform semantics-based trading strategies on prediction markets. One such strategy yields 14.12% net ROI after fees in a two-month period in 2026. Overall, we demonstrate the potential for agentic AI as a structural discovery layer for prediction markets.
Figures & tables
AType
Field
Description
Market
event_title
Title of the event associated with the market
market_info
Market question and resolution criteria provided as input
MarketRelation
mkt_info_leader
Exact input text for the antecedent market
mkt_info_follower
Exact input text for the consequent market
direction
Boolean relation sign: true for the same outcome and false for opposite outcomes
strength
Relationship strength: high, medium, or low
Table 1. Agentics input and output ATypes and their target fields used in the workflow.
Relation type
Leader proposition
Follower proposition
Agent rationale
Logical, +
Trump is impeached and convicted by the Senate before Jan. 20, 2029
Trump leaves office before Jan. 20, 2029
Conviction removes the President, so leader-Yes implies follower-Yes.
Economic, +
Fed target upper bound exceeds 3.5% after the March meeting
Three-month Treasury par yield exceeds 3.5% at quarter end
The short yield is strongly influenced by the policy-rate target.
Exclusion, −
Crockett wins the Democratic primary by at least 9%
The general election is Talarico versus Paxton
Crockett’s nomination rules out a matchup requiring Talarico as nominee.
Table 2. Abbreviated high-strength Agentic relations from the March output. Questions are shortened only for presentation; endpoint resolution uses the full exact text.
Metric
NLI
AAI
Difference/test
Resolution (all)
40.6%
62.8%
+22.2 [14.0,36.9] pp
Shared resolved
25.7%
57.0%
block p=.0018
Positive only
60.6%
66.7%
+6.2 [ − 10.6,17.3] pp
Negative only
40.5%
39.0%
− 1.5 [ − 14.3,16.8] pp
Frustration L/∣E∣
5.88–16.16%
0.324%
NLI bounded; AAI exact
Table 3. Latest-snapshot pooled outcome tests and February–March graph diagnostics. Intervals are cluster-bootstrap 95% CIs for AAI minus NLI.
Parameter
Description
Value
gˉh
Minimum edge for high-strength relations
0.75
gˉm
Minimum edge for medium-strength relations
0.60
Bh
High-strength base notional
$800
Bm
Medium-strength base notional
$50
pˉ
Minimum eligible contract price
$0.03
C
Maximum exposure per cluster
$2,500
Table 4. Parameters of the semantic trading strategy.
Metric
Value
Metric
Value
Net PnL
$1,412.21
ROI
14.12%
Sharpe
4.46
Max drawdown
− 33.56%
Trades
46
Win rate
60.87%
Invested
$3,121.89
Fees
$190.44
Avg. PnL/trade
$30.70
Avg. hold
12.5 days
Table 5. Results of an investable performance estimate. Sharpe uses daily realized PnL divided by initial capital.
Can Large Language Models (AI agents) aggregate dispersed private information through trading and reason about the knowledge of others by observing price movements? We conduct a controlled experiment where AI agents trade in a prediction market after receiving private signals, measuring information aggregation by the log error of the last price. We find that although the median market is effective at aggregating information in the easy information structures, increasing the complexity has a significant and negative impact, suggesting that AI agents may suffer from similar limitations as humans when reasoning about others. Consistent with our theoretical predictions, information aggregation remains unaffected by allowing cheap talk communication, changing the duration of the market or initial price, and strategic prompting, thus demonstrating that prediction markets are robust. We establish that "smarter" AI agents perform better at aggregation and they are more profitable. Surprisingly, giving them feedback about past performance has no impact on aggregation.
Prediction markets aggregate collective intelligence to forecast uncertain events, but their utility depends on reliable outcome resolution. Existing oracle systems tradeoff fast but brittle automation against accurate but costly human arbitration. Single-LLM oracles achieve meaningful accuracy but inherit all failure modes of their underlying model with no self-correction mechanism. We evaluate whether multi-agent LLM architectures can improve oracle resolution accuracy over single-model baselines. We compare independent aggregation and deliberative consensus against single-LLM baselines (GPT-5 Nano, DeepSeek V3, and Llama-3.3-70B) on 1,189 resolved prediction market questions from KalshiBench. All agents share a common evidence layer through Exa, with retrieval filtered by publication date to isolate reasoning from retrieval quality. Independent aggregation with confidence-weighted voting achieves the highest accuracy at 83.43 percent, outperforming the best individual model by 1.01 percentage points. Deliberative consensus degrades accuracy to approximately 76 percent, below every single-model baseline, attributed to error propagation during debate where confidently wrong models flip correct ones. Error correlations across models (0.529-0.689) explain why aggregation gains fall short of the theoretical Condorcet ceiling, placing a fundamental limit on ensemble approaches. Many questions resist correction by any multi-agent architecture, motivating escalation to human arbitration. We propose routing criteria for hybrid AI-human oracle systems: auto-resolving only unanimous, high-confidence questions yields 97.87 percent accuracy on 47 percent of the dataset, with inter-agent disagreement flagging the remainder for human review.
Forecasting future events has attracted growing attention as a testbed for general-purpose AI. A natural way to ground this evaluation is let the models trade in the prediction markets. Trading, however, requires more than forecasting. Moreover, recent benchmarks report a substantial gap between calibrated probability scores and the trading results. We propose Raven-Agent, to the best of our knowledge, the first autonomous trading agent for prediction markets. On a controlled replay over an archived decision set, our architecture achieves the only positive return and the only positive risk-adjusted return among all tested policies. We have released our code in https://github.com/Alchemist-X/predict-raven .
Yishu Wang, Yuxuan Wang, Jiaqi Deng +1
Hong Kong University of Science and Technology · Peking University · The University of Hong Kong +1