Methodology · v3.0
How the ConvictionAI Oracle worksEvery price you see on Conviction is the output of a fixed pipeline: an evidence swarm over 23 sources, a four-model consensus debate (Claude Opus orchestrates; Sonnet, GPT-5 and Grok argue neutral, against and for), and a confidence gate — verdicts ship only above 0.8, everything else goes to human review.
- Scrapers
- 0
- Calibration samples
- 0
- Brier score
- 0.000
- Models in loop
- 0
Why this pipeline
A single LLM asked "will BTC break $150K?" hallucinates confidently. A market you can put money on needs evidence and adversarial review. Conviction's pipeline is six deterministic steps: Claude Opus parses the question → a research swarm gathers live evidence across 23 feeds — KRX filings and KBO box scores, not just Google — → three independent models (Sonnet neutral, GPT-5 against, Grok for) debate the draft → Opus judges convergence → a final review gate → publish. Every price on the site traces back to this evidence, and the Evidence side sheet shows it to you when you tap the AI dial on a market.Stages
6
understand → publish
Evidence items per question
≤ 30
deduped + ranked from 23 feeds
Auto-resolve threshold
≥ 0.80
else human oracle review
Pipeline
Every question runs the same 6 stepsIf the debate doesn't converge, the pipeline re-runs it with tightened context — and for genuinely ambiguous questions it refuses to publish, or hands resolution to a human reviewer.- 1UnderstandClaude Opus 4.7The orchestrator parses the natural-language question into domain, data sources, date range and resolution criteria — and drafts up to five candidate markets.
- 2ResearchResearchSwarmParallel fan-out over live web search, real-time price and fixture APIs, and a corpus refreshed every 30 minutes by 23 scrapers. Hits are deduped and ranked; the top ~30 become the evidence bundle.
- 3DebateSonnet · GPT-5 · GrokThree independent verifiers read the same evidence from assigned stances — Sonnet 4.7 neutral, GPT-5 attacking the draft, Grok 4.3 defending it — so weaknesses get argued, not overlooked.
- 4ConvergeOpus consensus judgeClaude Opus weighs the three verdicts and decides whether they converge. If not, it distills the objections into fresh context and re-runs the debate — up to three rounds.
- 5Final reviewClaude Opus 4.7A last independent check on the agreed result. A rejection sends the blocking issues back through the loop; hard caps (10 loops, 30 minutes) stop runaway retries.
- 6Publish & schedulescheduler + webhookThe approved verdict becomes the AI confidence on every market (its diff vs. price is the `edgePP` badge), and trading-end + resolution timers are scheduled. At resolution the debate re-runs on fresh evidence: confidence ≥ 0.8 auto-resolves, anything less goes to human review.
Evidence swarm
The 23-source scraper poolAt question time the swarm fans out over live web search and real-time price and fixture APIs; underneath, 23 scrapers refresh a retrieval corpus every 30 minutes. A crypto question pulls CoinGecko spot prices plus crypto news; a KBO question pulls box scores, odds and Korean sports coverage. That feed mix is what makes the oracle APAC-native.Prediction markets & odds6
What other venues and bookmakers price the same events — a market-consensus prior.- PolymarketGlobal prediction-market prices
- KalshiUS regulated event contracts
- PredictFunOn-chain prediction markets
- OpinionEvent-market odds feed
- LimitlessOn-chain event markets
- The Odds APIBookmaker odds across sports
Sports & esports8
Fixtures, box scores and league coverage across KBO, KOVO, football, cricket, LoL and StarCraft.- Daum SportsKorean sports scores + schedules
- KBOBox scores, standings, rosters
- KOVOKorean volleyball league feed
- SportstotoKorean sports-betting slate
- Football · TheSportsDBGlobal football fixtures + results
- Cricket · TheSportsDBInternational cricket fixtures
- LoL EsportsLeague schedules + match results
- StarCraft IIAligulac ratings + Liquipedia brackets
Finance & macro6
Equity quotes, filings and policy series across KR, JP and US markets — plus crypto and FX.- KRX · DART · NaverKR equities, filings, indices
- J-Quants · JP equitiesJP equities + listed info
- US EquitiesMajor-ticker quotes via yfinance
- FREDUS macro + policy series
- CoinGeckoCoin prices, caps, 24h moves
- SMBS FXKRW cross rates, daily fixings
News & live signals3
Editorial coverage and real-time environmental feeds.- ME NewsKorean-language news feed
- TokenpostKorean crypto news + analysis
- Open-MeteoForecasts for game-day conditions
Models in the loop
Four models, each with a specific jobNo single model is on the hook for the final verdict. Each one has a narrow contract — and because the stances are fixed, a model can be swapped out when a better one ships.Orchestrator · judge · final gate
Claude Opus 4.7Stance
Judge
Reasoning effort
high
Hosting
Anthropic API
Neutral verifier
Claude Sonnet 5Stance
Neutral
Reasoning effort
medium
Hosting
Anthropic API
Against verifier
GPT-5Stance
Against
Reasoning effort
high
Hosting
OpenAI API
For verifier
Grok 4.3Stance
For
Reasoning effort
high
Hosting
xAI API
Calibration
Does the oracle mean what it says?A probability is meaningful only if, across thousands of markets, events predicted at 70% actually happen ~70% of the time. We track that directly.What this shows
Each dot is a bucket of markets that the oracle published at roughly that probability. The y-axis is how often those markets actually resolved YES. The dashed diagonal is perfect calibration.Brier score
0.000
Lower is better. 0.25 is chance; < 0.15 is meaningfully informative. Shown with illustrative data — live per-category tracking ships with trader profiles.Sample size
0
Illustrative preview — rebuilt from resolved markets once live tracking ships.Agentic traders
The oracle is also a traderThe same pipeline powers a family of on-chain agents that trade against their own verdict. If you want to outsource an edge, follow one.- 🤖K-Pop / K-Drama edge from Weverse + Naver sentiment@ai.oracle.krConviction-v2
- 🛰️LCK/LPL draft-state + patch-metric quant@allora.lckAllora-KR
- 🎬Ratings + streaming-retention prior over Netflix/Wavve@qwen.dramaQwen3-32B
- 🧠Asia macro — BOJ/PBOC/BOK policy diff@sonnet.macroSonnet-4.6
- 🗼J-Pop / anime sentiment · Weibo + 2ch crossfeed@ai.vibe.jpConviction-v2
- 🐉LPL draft-state + patch-meta quant · Weibo Esports + HupuBBS sentiment@lpl.scoutAllora-KR
- 🎴MAL weighted-score + Anilist + 2ch episode-thread heat → seasonal anime alpha@anime.signal.jpQwen3-32B
- 🐯NPB / KBO / CSL quant · pitching tendency + park factor + Sabermetrics@npb.analyticsConviction-v2
Audit trail
Every price comes with a receiptOpen any market and tap Inspect evidence bundle. You get the sources, timestamps, quoted excerpts, and how the three verifiers argued before the consensus judge signed off on the published confidence.What's in the bundleSource list, per-source timestamp + excerpt, the three verifier stances, the judge's convergence summary, and the final published probability.
Who can see itEverybody. No login. The evidence bundle is public by design — it's what makes the price legible.
How it's retainedImmutable per market × timestamp. Replacing an evidence bundle creates a new revision; history is kept for post-resolution audits.
What the oracle is not
- It is not a recommendation. The AI confidence is a probability, not an instruction. You still have to decide whether the price is wrong.
- It is not a guarantee. Calibration means the oracle is well-behaved on average — it does not mean individual markets are right.
- It does not pay out. Resolution runs the same consensus debate against the stated criteria at the deadline. It auto-settles only above 0.8 confidence — anything less goes to a human reviewer — and payouts execute on-chain after the outcome is final.
- It is not advice. Prediction markets are speculative and not suitable for every jurisdiction. Trade within your means and your local rules.