Use cases
Concrete things this dataset is built to make possible — grounded in
what's actually in v1.0 today (5 cities, daily observations
back to 1869, forecasts/errors/revisions/probability curves from 2021
onward). See the explorer for live numbers.
Point-in-time backtesting, no look-ahead bias
Every forecast row records its own forecast_issued_at
and model_run_timestamp. Filter to
forecast_issued_at <= decision_time and you get a
provably point-in-time-correct feature set — no manual
reconstruction required.
Forecast-provider skill evaluation
forecast_errors pairs every forecast with its eventual
observed outcome and a signed error, already point-in-time safe —
slice by city, season, or lead time.
Forecast revision analysis
forecast_revisions lines up the 24h/48h/72h-ahead
forecasts for the same target date side by side, plus the deltas
between them — the raw material for "how much does the forecast
move before expiry" research.
Calibration research
probability_curves ships an out-of-sample-calibrated
P(observed > forecast + k) for a standardized offset grid,
computed from strictly-prior-date training errors only — a ready
baseline to benchmark your own calibration approach against.
Weather-sensitive strategy research
Daily max temperature plus forecast error/revision history by city is a standard input to degree-day demand models (energy), crop- stress models (agriculture), and parametric-insurance research.
What this dataset deliberately does not include (yet)
- Kalshi market data (contracts, quotes, settlement) — redistribution rights still unresolved.
- A commercial redistribution license for Open-Meteo-derived fields — today's free tier is non-commercial.
- Intraday/sub-daily granularity, or cities beyond the initial 5.