Prediction market arbitrage sounds like free money: buy YES on one site and NO on another, pay less than a dollar, and collect a dollar whatever happens. I built a scanner to find these trades across 131,000 markets. The first scan found a $5,915 "opportunity". It was two different bets wearing the same title.
This is the story of building that scanner in two days, what the real numbers look like, and why I ended up handing the hardest job to ChatGPT instead of doing it myself.
The Idea
Prediction markets like Polymarket and Kalshi let you buy a contract that pays $1 if something happens and $0 if it does not. The price is the crowd's probability. If YES costs 40 cents, the market thinks it is roughly a 40% chance.
The same event is often listed on several venues, and different crowds price it differently. If one venue sells YES at $0.40 and another sells NO at $0.55, you can buy both for $0.95. One of them must pay $1. You never need to predict anything. You just need the combined price to be under a dollar.
Finding these by hand means opening dozens of tabs and comparing prices. So the idea was simple: one button, Run Scan, and a ranked list of every trade like that, biggest first. I place the trades myself; the tool only looks.
Why It Is Not Free Money
Before writing any code I listed the ways a "risk-free" trade quietly stops being risk-free:
- The wording — two markets with the same title can resolve on different rules, deadlines or sources. Then it is not a hedge, it is two bets.
- The liquidity — a 5% spread is meaningless if only three contracts are for sale at that price.
- The fees — both venues take a cut, and on thin edges the fees are the whole profit.
- The timing — the gap can close while you are placing the second order.
I thought the first one would be a footnote. It turned out to be the whole project.
What I Actually Built
A local web app I run on my own PC. No API keys for the venues, no wallet, no order placement — it only reads public prices.
- Data — the open-source ccxt library, which recently added a
predictionmodule covering nine venues with one interface. - Matcher — finds the same question on two venues. It extracts numbers, dates, years, "above/below", names and negations, and rejects any pair where those disagree.
- Pricing — walks both order books level by level, adding contracts while the pair still costs under $1 after fees, capped at $500 per trade.
- App — Python and FastAPI, SQLite, and a plain HTML page with a big Run Scan button. One scan takes about two minutes.
Step 1: Test the venues before trusting them
The first thing I did was probe all nine venues with no API keys and write down what actually came back. Three surprises came out of that half day:
- Two of the nine venues refuse to show even public prices without an account key.
- Kalshi's standard market list returned 1,000 markets, and all 1,000 were multi-leg sports parlays — useless for matching. Its raw events endpoint returned 85,000 real markets in 26 seconds, with the full rules text included.
- The library reported Polymarket's fee as a flat 10%. Polymarket actually charges a small, price-dependent taker fee that is different in every category. Had I trusted the number, every trade would have looked unprofitable.
None of this was in any documentation. It only showed up because I looked at real responses before building on top of them.
Step 2: Match questions, carefully
Matching is the hard part. Kalshi alone lists over 100,000 open markets, many of them nearly identical: "Bitcoin above $120,000 on Oct 3", "above $121,000", "above $122,000". A fuzzy text match would happily pair the wrong ones.
So the matcher is deliberately suspicious. Thresholds must be identical. Dates and years must agree. "Above" can never match "below". If each side names someone the other never mentions, the pair is rejected. Head-to-head sports markets have to name the same winner.
Step 3: Price it like a real trade
The listing price is only the best offer. The scanner fetches the full order book for both legs and fills them together — 120 contracts at 40 cents, then 80 more at 41 cents — stopping the moment the next pair would cost a dollar or more, or when it hits my capital cap. Then it subtracts each venue's real fee formula. What comes out is a size, a cost, and a net profit I could actually get.
The First Scan Found $5,915. It Was Fake.
The first full scan pulled 128,404 markets, matched 731 pairs, and priced the top 150 in under two minutes. The top of the list looked incredible.
#2 "Will Spain vs. Czechia end in a draw?" + "Spain vs Czechia: Spain wins" — $2,422 profit
#3 "Will Drake be the top artist in the US?" + "Who will be the top Spotify artist this year? Drake" — $1,108 profit
#4 "Team Spirit vs Team Yandex" + "Team Yandex vs. Team Spirit: Team Yandex wins" — 31% return Every one of these is a different bet, not a hedge. #4 is the scariest: buying YES on the first and NO on the second pays only if Team Spirit wins. Both legs lose if Yandex wins. It is the same bet placed twice.
The titles were close enough to fool a text matcher. The rules were not close at all. "Trump meets Putin by October 31" on Polymarket means in person. The Kalshi version says "including phone calls". One phone call and both legs of my "risk-free" trade lose.
// The uncomfortable bit
The bigger the spread, the more likely it is a mistake. Markets are watched by a lot of people with money on the line. A real gap between two identical contracts lasts minutes and is a few cents wide. A 30% gap that sits there all day is almost always telling you the two contracts are not the same thing. The scanner's most exciting results were its least trustworthy ones, and I had sorted the list to put them on top.
So I added more rules: reject pairs where both sides name different subjects, catch negated wording ("no Gemini release" vs "Gemini release"), check who wins in head-to-head markets, and flag anything with a spread above 10% as suspicious.
I Could Not Verify 726 Matches. Nobody Could.
The next version had three tabs: Opportunities, Needs Review, and All Matches. I opened All Matches, clicked a pair — "Will Donald Trump attend UFC 332?" — and the panel told me to "Buy 0 NO on Limitless" and "Buy 0 YES on Kalshi".
It was technically correct. Covering both outcomes cost $1.03, so the right trade size was zero. But it was a terrible screen. I was looking at 726 pairs, most of them with no gap at all, and the tool expected me to read the rules of each one and decide if they matched. That is not a tool. That is homework.
Your judgement is still needed for the few pairs that are profitable but uncertain. For everything else, the honest answer was to let a model read the rules.
Letting ChatGPT Read the Small Print
I wrote a checklist of six things that have to agree before a pair is a real hedge, and asked ChatGPT to check them for every pair with a real price gap:
- Same subject and numbers
- YES/NO direction — does one leg pay exactly when the other loses?
- Same deadline and time zone
- Same definition and resolution source
- Same rules for cancellation, postponement and ties
- Price and payout in every scenario
It answers with a strict JSON verdict — "hedge holds", "risky" or "not a hedge" — plus a table of outcomes showing what each leg pays. Pairs with no gap are never sent, and every answer is cached against the exact contract text, so a pair is only ever paid for once.
Picking the model with real cases, not benchmarks
Instead of guessing which model was "best", I took 16 pairs from my own scans where I knew the right answer — the Team Spirit trap, the phone-call meetings, the Drake mismatch, and genuine matches — and ran each candidate model on all of them, twice.
| Model | Caught every trap? | Cost / pair | Speed |
|---|---|---|---|
| gpt-5.6-luna, medium reasoning | Yes, in both runs | $0.002 | 15s |
| gpt-5.6-luna, low reasoning | Called the worst trap only "risky" | $0.0017 | 11s |
| gpt-6-luna, medium | Let one trap slip once | $0.001 | 18s |
| gpt-6.1-sol, low | Let one trap slip | $0.013 | 22s |
The most expensive model was not the best. A mid-priced small model with a bit more thinking time was. No model ever called a trap "safe", which was the property I cared about most.
The AI was right and I was wrong
Two of my "genuine match" labels kept getting marked as not a hedge. I assumed the models were being fussy, so I went and read the rules myself.
// DeepSeek IPO before 2027
Polymarket needs the IPO to be completed. Kalshi pays YES once the IPO is priced or gets a ticker. If pricing happens in December and trading starts in January, both of my legs lose.
// US confirms aliens before 2027
Kalshi defines "the Cabinet" to include the White House Chief of Staff, the CIA Director and others. If one of them made the statement, Kalshi could resolve YES while Polymarket resolves NO.
Both were real hidden differences that I had missed while writing the test cases. The tests were supposed to grade the model, and the model ended up correcting my answer key.
What It Looks Like Now
The final version has two tabs and no noise. A pair only appears if it has a real gap, at least 20 contracts available, at least $5 and 1% net profit, and the AI has not found it to be a fake hedge. Everything else is counted in one small "filtered out for you" card and never shown.
The latest scan: 131,499 markets, 665 matched pairs, 127 with any price gap at all. ChatGPT removed 49 as not real hedges, 76 were too small, and two were tradable:
| Pair | Net profit | Return | Size | AI verdict |
|---|---|---|---|---|
| Letitia James arrested before 2027 — YES on Polymarket, NO on Kalshi | $7.22 | 1.86% | 396 | Risky |
| Adam Schiff arrested before 2027 — YES on Polymarket, NO on Kalshi | $6.08 | 1.22% | 506 | Risky |
Even these come with a warning: Kalshi counts surrendering on an indictment without an arrest warrant as an arrest, and Polymarket's list may not. That is a narrow case, and the AI's job is to tell me exactly which sentence to go and read.
"Real arbitrage looks boring: one or two percent, three months away, a few hundred contracts. Anything more exciting deserves suspicion."
// What 131,000 markets taught meWhat I'd Tell Myself on Day One
The spread is a warning label, not a prize
I sorted by profit and put the biggest mistakes at the top. In a market full of people with money on the line, a large gap that nobody has taken usually means you have misunderstood something. Treat outliers as bugs until proven otherwise.
Look at the data before designing around it
Half a day spent calling each venue and reading the raw responses saved me from a fee that was off by 10x and a market list made entirely of parlays. Documentation tells you what a system is supposed to return. Only a real request tells you what it does return.
Don't ask a human to do what a machine can sort
A list of 726 matches felt thorough. It was actually me pushing the work back onto myself. A good tool shows you the few decisions that need you and quietly handles the rest.
Choose models with your own examples
Sixteen labelled cases from real data told me more than any leaderboard. They showed which model caught the dangerous traps, what it cost, and — unexpectedly — where my own labels were wrong.
Count the two kinds of mistakes separately
Missing a real 1% trade costs me $5. Trusting a fake hedge can cost me $500. When the two errors cost such different amounts, the tool should lean hard toward caution, even if that means fewer results.
Was It Worth It?
Yes, but not because it found free money. It mostly found that free money is rare, small and slow, which is exactly what you would expect if markets mostly work.
What I got instead is a tool that does the boring part well: it checks 131,000 markets in two minutes, prices them honestly, and has an AI read the small print, all for about a fifth of a cent per check. When a real gap does open, I will see it, with the rules already compared and the order book already sized.
It is read-only, it never places a trade, and nothing here is financial advice. It is a very fast, very skeptical research assistant — which, it turns out, is what an arbitrage scanner should be.
Back to Blog