Model Performance — Live
The scoreboard.
Every EVIO model probability is stored next to the market's price at the same instant and scored after settlement — hits and misses alike. Below: every model we run and how much evidence each one actually has, the NFL model against the consensus line, the moneyline and the total, and the live head-to-head ledger across every domain.
Every model we run
Four lanes, one rule: the score is published.
Each lane below is a live model priced against a real venue contract. The state on each tile is how much evidence exists today — a backtest is not a live record, and a lane that has captured markets but settled none has nothing to brag about yet. We label the difference instead of blending it.
Figures refresh from /v1/markets/performance and /v1/nfl/performance — both free and unauthenticated, so anyone can reproduce this table.
Live head-to-head ledger
EVIO probability vs. market price
Settled venue contracts only — one final pre-close call per market, scored by the venue's own result. Brier: lower is better.
Methodology & honesty rules ▾
Whenever an EVIO model holds a probability for a live venue contract, the venue's public prices and our probability are snapshotted in the same row at the same instant.
After settlement the venue's own result is written back. Aggregates use the last pre-close capture per market — one final call per contract, misses included. Backtest tickets are flat 1 unit: spread and total at −110, moneyline at the actual quoted odds, so the vig is paid everywhere a real ticket would pay it.
The moneyline view exists to publish an uncomfortable number on purpose: our win probabilities are calibrated (train and holdout Brier match to the third decimal) but the market's are sharper — so we sell them as pricing inputs and do not recommend moneyline picks. If we hid that, nothing else on this page would be worth believing.
Model research, not investment advice. Hypothetical results at captured public prices; EVIO places no orders and holds no customer funds. Event markets involve risk. Raw data: /v1/markets/performance · /v1/nfl/performance
How each model works
Disclosed, not black-boxed.
A prediction you cannot inspect is a prediction you cannot price. Every model we sell is specified here in full — the inputs, the math, and the contract it settles against.
Weather — daily settlement temperature
wx_sigma_v0A per-degree distribution over the official daily maximum or minimum at a named settlement station, out to six days. The forecast mean comes from the NWS gridpoint product; the spread is a horizon-dependent sigma fitted to our own archive of settled climate days. The distribution — not a point forecast — is the product.
- Settles against
- Kalshi KXHIGH* / KXLOW* daily temperature ladders, and the equivalent Polymarket contracts.
- Why we have an edge here
- We store the exact instruments those contracts settle on — 9 stations, 72,846 station-days — and read the climate day in local standard time, the way the venue resolves it, not in UTC.
- Known gap
- The in-day blend. Once part of the day has already happened, the market knows the realized high and our remaining-horizon forecast does not. We suppress in-day captures after local noon rather than publish a number we know is stale.
MLB — game win probability
mlb_log5_v0Team strength from Pythagorean expectation on run differential (exponent 1.83), shrunk toward .500 early in the season, combined through log5 for the matchup, then adjusted for home field and the announced starting pitcher's ERA.
- Settles against
- Kalshi KXMLBGAME per-game moneyline contracts, matched by discovery rather than constructed — the ticker embeds first-pitch time and doubleheader game number, so guessing it is how you price the wrong game.
- Adjustments
- Home edge 3.5 points of win probability; starter ERA worth 2.0 points per earned run, capped at ±6 points, and skipped entirely when the starter is unannounced or under 20 innings.
- Honest limits
- No bullpen model, no park factors, no lineup or platoon handling, no travel or rest. It is a strong baseline, not a finished product, and the ledger will say so.
Crypto — strike-grid price distribution
crypto_rv_v0A zero-drift lognormal over BTC and ETH, with volatility measured from the realized variance of the last 120 one-minute closes and scaled by the square root of the horizon. Deliberately driftless: we price dispersion, and we do not claim to know direction.
- Settles against
- Kalshi KXBTCD / KXETHD hourly ladders, which resolve on a 60-second average of the CF Benchmarks real-time index.
- Volatility bounds
- Annualized sigma floored at 15% and capped at 250%, so a dead tape or a single violent minute cannot produce a confident-looking absurdity.
- Why the headline number is small
- Most strikes on an hourly ladder are far out of the money and both sides price them near zero. Those are free to get right. The contested subset — both probabilities between 5% and 95% — is the only number worth reading.
NFL — spread, total, win probability
nfl_v1Team form from exponentially-decayed per-play efficiency (EPA per pass and rush, CPOE, sacks, takeaways, first downs, penalty yards) over a 16-game window with an 8-game half-life, ridge-regressed to a point spread and a total, with a separately fitted logistic model for win probability.
- Settles against
- Consensus closing spreads and totals, and Kalshi KXNFLGAME moneyline contracts.
- Validation
- Configured entirely on walk-forward cross-validation inside the training era; the 2024–25 holdout was scored once, after the configuration was frozen.
- Honest limits
- The market is sharper than we are on every measure: spread error 10.15 vs 9.69, win-probability Brier 0.221 vs 0.206. Injuries, line movement and QB-specific ratings are not in the model yet. Details in the tabs above.
