Measuring noise in AI models by forecasting the top Polymarket events.
This benchmark estimates each model's judgment noise by comparing the consistency of its forecasts for highly impactful markets across many topics. Understanding noise in judgment helps us understand the degree to which a model's forecasts and opinions are inconsistent. Achieving 0% noise is possible but doesn't mean the forecasts will be perfectly accurate, only that they will be internally consistent.
Each model produces both individual forecasts for the odds of each market, eg "Will JD Vance win the 2028 US Presidential Election?" both for the "Yes" and "No" outcomes, multiple times. Then the models judge which of two separate market outcomes is more likely, eg "Billionaire one-time wealth tax passes in California election 2026?" vs "Will Marine Le Pen win the 2027 French presidential election?", for the 4 different outcome combinations of each pair of markets.
Loading…
A single score per model that combines five noise measures with a simple average.
An answer is one call to the model, counting both the probability prompts and the head-to-head ones, since the composite is built from both.