CALIBRATION
HOW THE MODEL CHANGES
ATLAS does not retune itself. Refitting a model on a few dozen outcomes fits noise, and a system that quietly adjusts its own weights cannot be audited by the people relying on it. Instead every candidate it prices is kept — including the ones it never published — and graded as its market resolves, with the parameters changing by hand, on that evidence, and each change recorded below.
The evaluated field
Everything priced, not just what was published.
Published picks are chosen on the largest gaps between the model and the market, and the largest of many noisy estimates is disproportionately the one whose error pointed our way. A record made only of picks flatters the model by construction. These are all the candidates, graded the same way.
Observations
828
Distinct games
226
Settled
381-336
Awaiting result
111
Does conviction predict outcome?
Every band of model edge, and how it did.
Bands are the model’s probability minus the price paid, in percentage points. Win rate is shown because people expect it, but it is descriptive only and is not comparable between bands — each mixes short prices with long ones, so a correct model can still show a falling win rate. Brier score is the price-aware measure, lower is better, and the market’s is computed over the same rows as the model’s.
| Band | Settled | W–L | Win % | Model Brier | Market Brier | Sample |
|---|---|---|---|---|---|---|
| Model below price (gap < 0) | 165 | 88–77 | 53.3% | 0.248 | 0.2502 | counts |
| 0–4 pts above price | 126 | 81–45 | 64.3% | 0.2327 | 0.2343 | counts |
| 4–8 pts above price | 119 | 77–42 | 64.7% | 0.2183 | 0.2295 | counts |
| 8–12 pts above price | 60 | 28–32 | 46.7% | 0.2611 | 0.2406 | counts |
| 12+ pts above price | 22 | 10–12 | 45.5% | 0.2752 | 0.2394 | under 25 |
No verdict
492 graded results so far — too few for this table to support a verdict about which band is better. Read it as a description of what has happened, not as evidence about what will.
Why it takes so many ▸Why it takes so many ▾
Ranking five bands by win rate is not a test. Under a model with no skill the bands land in a tidy ascending order surprisingly often, purely by chance. Telling a real 5-point win-rate difference from luck needs on the order of 1,500 graded results per band — and results inside one game move together, so three lines on the same blowout count for far less than three independent facts. The stronger measure available sooner is the paired model-versus-market Brier difference over the whole field, and that is not being claimed yet either.
These are candidates ATLAS priced, not bets it placed. No money was staked on any of them, so nothing here is a return. The bands group candidates by how far the model's probability sat above the price. Win rate cannot be compared between bands — each one mixes cheap prices with expensive ones, and a cheap winner and an expensive winner are not the same achievement. That is what the Brier columns are for: they score how close each probability was to what actually happened, so lower is better and the two columns are directly comparable. One thing that favours us: the market's Brier is scored against the price you would have paid rather than the midpoint, which makes the market look slightly worse than it was. Every candidate here was recorded before its game started, with the model version that priced it. Nothing was fitted afterwards.
Change log
What has actually changed, and what forced it.
Every entry is a deliberate edit with a reason, verifiable in the commit history. No parameter has ever changed automatically.
Publish rule bounded in runs, not expected value
2026-08-17
An EV threshold is a threshold on disagreement, and near the centre of the run distribution a quarter-run moves the probability hard while in the tails a full run barely moves it. Two calls priced above 10 points of EV disagreed by 0.67 and 0.27 runs. Bounding runs bounds what the model actually claimed; any disagreement of a run or more is now refused, whoever moved.
Publishing suspended entirely for a day
2026-08-16
On disagreements of a run or more the model sat 0.45 from the league average while the market sat 1.31 away, and the model was the one nearer the average in 130 of 149 cases. The old threshold was selecting for the model failing to move, priced as conviction. Every market was withdrawn and the reasoning published before anything resumed.
Negative Binomial dispersion refit, 15 → 10
2026-08-11
Fitted r ≈ 9.7 against seven Polymarket price ladders. The old value made the run distribution too tight, understating tail outcomes.
Park factor marked provisional, blocking affected picks
2026-08-11
Sutter Health Park had no established factor. A pick built on the guess was published, and it lost — the market was right.
Closing prices quarantined, CLV suppressed
2026-08-11
Every stored closing price failed a plausibility guard. A CLV computed from them was meaningless, so no number is shown rather than a wrong one.
College football confidence capped at 35 on prior-season data
2026-08-11
A completed prior season carries ~12 games, so week 1 scored maximum confidence off a roster that no longer exists.
When a parameter is allowed to move
The threshold is set before the data is read.
A band needs at least 25 settled observations before it is treated as evidence. Below that its numbers are shown and no conclusion is drawn from them, because a five-observation bucket at 80% means nothing. Deciding the bar after seeing the results is how you talk yourself into a change the data does not support.