June 1, 2026 baseline

Minnesota Population Explorer

More
Ask panel
Loading local data summaries...

Explore Population

Denominator: loading

Synthetic Persona Profile

Summary from stored attributes.

Search for a permanent synthetic ID or choose a random registered persona within the current filters.

Demographic Group Comparison

Mixed-data groups are the default; numerical k-means is available as a benchmark.

County Composition

Stability Checks

Historical Scenarios And Official Results

Simulated histories describe the current synthetic cohort under assumptions; they do not reconstruct original electorates.

Simulated Ballot Categories

Official County Actuals

Official results are aggregate county observations. Scenario selection changes simulated histories above, not these official totals.

Methods And Data Quality

Official aggregate observations, statistically simulated attributes, and future fictional characterization are kept distinct.

About the Explorer & Election Forecasts

How the models work, where the evidence comes from, and what the results can—and cannot—tell you.

Two separate kinds of modeling

The population explorer describes a supplied synthetic Minnesota population with a June 1, 2026 baseline. Its people and individual historical ballots are simulated, not observed real voters. Official aggregate election results are kept separate from those simulated histories.

The 2026 election forecasts are statewide and county aggregate estimates for Governor/Lieutenant Governor, U.S. Senate, Attorney General, Secretary of State, and State Auditor. The new Synthetic-population + AI model uses the registered baseline as a counted population; it does not simulate a ballot for each person, and the Explore county and demographic filters do not affect a statewide run. Judicial, legislative, and local races are outside this model.

Synthetic-population + AI model

The default model is OpenAI GPT-6 Astra, with the exact API identifier gpt-6-astra. The run form can select another available model. Each saved run records the requested model and the actual identifiers returned for research and estimation; check Models Executed rather than assuming a run used the default. Unexpected model substitutions are rejected, not silently replaced.

  1. Deep search: GPT-6 Astra uses the Responses API and web search to review candidates, historical results, polling, statewide fundamentals, and accessible prediction-market evidence, guided by the selected as-of date and research notes.
  2. Aggregate reasoning: research is summarized into demographic- and county-level reasoning, then explicit probability-scale turnout/support parameters and an uncertainty note. AI reasoning drives these parameters; it is not a posthoc blend of a separate top-line forecast.
  3. Population calculation: the engine applies those parameters to exact counts for all 3.866 million registered baseline records (3,866,086 in the supplied snapshot), represented by exchangeable demographic/county cells. It makes no per-person AI guesses, does not infer true political preferences, and does not mutate the baseline.
  4. Correlated draws and outputs: 5,000 shared-state/county-correlated draws produce statewide turnout and D/R/Other results plus county outcomes. The full population count, baseline date, fingerprint, cells, parameters, and uncertainty note are saved with the snapshot.
  5. Validation and saving: the app checks numeric ranges, consistency, source URLs, population counts, and actual model identity. Incomplete or invalid output fails rather than becoming a forecast. Research, citations, assumptions, outputs, and the separate benchmark are saved together.

Primary turnout comes from AI group probabilities and registered-population counts. The entered Total Votes on an AI run is visibly BENCHMARK-ONLY; it is not substituted for the AI turnout. Other share remains a fixed user assumption; each draw rounds its statewide Other target to the nearest integer vote, so the count can differ from that percentage target by at most one vote. New AI runs perform new research, while completed research can be reused if estimation needs to resume. Web search does not guarantee complete coverage or factual correctness, and validating a citation URL does not independently verify claims. An earlier as-of date is not a leakage-proof historical backtest: today's model may know later events.

Legacy AI-led runs and statistical benchmark

Older snapshots with results.mode = ai-led remain separate legacy AI-led estimates. They use supervised public-web research and a structured aggregate judgment with subjective intervals and win probabilities; they do not use the synthetic population to infer individual preferences. They are not relabeled as the new population model.

The benchmark runs 5,000 seeded Monte Carlo draws. It begins with the historical Democratic share of the two-party vote and a normal prior with a 5-percentage-point standard deviation. The same seed and inputs reproduce the statistical result; they do not make AI output deterministic.

Eligible polls are converted to two-party shares. Their influence declines exponentially with age (a 45-day decay scale) and increases with sample size using a capped square-root weight. Sampling uncertainty has a 3-point floor, combined polling precision is capped to account for common polling error, and the final polling-informed standard deviation cannot fall below 3 points. With no eligible polls, the model uses the historical prior alone.

Each draw splits the vote remaining after the user-specified Other share between the two major parties. Results show mean shares, 5th–95th percentile ranges, and the fraction of draws in which each major party wins. The entered total-vote assumption converts shares into vote counts; it is not a turnout forecast. Minor-party and write-in votes are grouped as Other, held fixed, and have no individual win probabilities.

The benchmark is not blended into either AI model. It is displayed alongside AI results as a reference. Differences reflect different assumptions and judgment, not proof that any approach is correct.

Election data and source citations

  • Historical priors: Minnesota Secretary of State 2022 general-election results for the four state offices, and 2020 general-election results for the same U.S. Senate seat. These are aggregate observations from different candidates and electorates—not 2026 support measurements.
  • Initial polling catalog: SurveyUSA/KSTP election poll, fielded August 13–17, 2026, released August 18, with 661 likely voters. The catalog includes one poll for each of the five races. Editable polls remain user-supplied inputs, not independently verified data.
  • Candidates: race-specific Ballotpedia links appear on each race card—for example, Governor/Lieutenant Governor and U.S. Senate—alongside relevant polling sources. Use the Minnesota Secretary of State for official election information. The catalog is not a certified complete ballot.
  • AI evidence: the actual retrieved sources vary by run. Open a saved run's evidence and full research report to inspect its citations, dates, rationale, and any market observations. There is no guaranteed live market feed or fixed list of research publishers.

Population data and geography

The explorer reads the imported population and supporting-person files through DuckDB; it does not fetch a live voter file. The registered baseline is dated June 1, 2026, and registration publication/live-table timing does not turn it into a current election-day file. The supplied source inventory documents Minnesota registration counts, 2020–2024 ACS five-year demographic tables, and ACS Public Use Microdata Sample documentation. Aggregate targets and survey donor records do not establish the traits or preferences of any identifiable voter.

The county map uses the U.S. Census Bureau 2024 county cartographic boundary file, 1:500,000, filtered to Minnesota. ACS estimates have sampling uncertainty and multi-year reference periods; they are not direct June 2026 measurements. More → Methods provides the imported dataset's validation and methodology notes.

How sources are updated

  • Population and maps: bundled snapshots, not automatically refreshed or aged to today's date. Updates require deliberately importing/rebuilding the source data and rerunning validation.
  • Race catalog: the initial candidate, prior, and poll snapshot was verified September 12, 2026. There is no scheduled automatic catalog refresh. A maintainer must review sources and update the catalog; the Races page displays its verification date and individual source links. New AI research does not rewrite that catalog.
  • Poll edits: changes in the run form affect that run's submitted inputs, not the curated source catalog. Review field dates, publication dates, sample sizes, and provenance before running.
  • AI research: requested when you initiate a run, rather than continuously in the background. Research timestamps, actual executed model IDs, and source citations are preserved with the run; coverage and publication timing may still lag real events. Research completed before an interruption may be reused for estimation.
  • Saved results: snapshots do not refresh as sources change. Run again to obtain a new dated estimate. After election day, actual results may be entered separately for comparison without overwriting the forecast. Actuals are user-entered and are not automatically verified, even if labeled certified.

Accuracy disclaimer

This is an academic experimental research tool, not an official election forecast, certified ballot, or guarantee of an outcome. Neither the Synthetic-population + AI model nor legacy AI-led model is fitted or predictively validated, and neither is empirically calibrated. A reported 78% win probability, for example, is not a demonstrated 78% real-world success rate, and no benefit of demographic modeling has been validated.

The synthetic model covers the registered baseline only: it makes no post-baseline or newly eligible population adjustment. It does not use fabricated persona ballots as calibration and does not turn demographic attributes into true inferred political preferences. Results can be wrong because of stale or incomplete sources, polling bias, turnout changes, correlated errors, candidate changes, unforeseen events, simplifying assumptions, or AI mistakes. Uncertainty ranges are conditional on the method and inputs and may omit important risks. The fixed Other share can understate third-party uncertainty. More draws do not correct bad assumptions.

Future actual results can support later model assessment and possible fitting, but comparison does not prove accuracy; fitting is not currently automatic. AI may misread sources or generate unsupported claims despite validation. Read the underlying sources, compare assumptions, and consult official election authorities. Do not use these estimates as the sole basis for voting, financial, betting, or other consequential decisions. Synthetic personas must not be interpreted as real individuals or evidence of their political beliefs.

Access and AI costs

AI calls use the owner's OpenAI account and incur charges there. This app currently has shared access without user authentication: anyone with access can view saved runs and initiate AI requests. The one-active-job and ten-requests-per-rolling-day limits are request limits, not a dollar spending cap. Do not include private information in run names or research notes.

2026 Statewide Races

Baseline data and forecasts for the upcoming state-wide elections. Note: Forecasts are statewide only and do not use the Explore county/demographic filters.

Loading races and forecast controls…

Saved Forecast Runs

Compare simulation models and actual results. Note: Forecasts are statewide only and do not use the Explore county/demographic filters.

Research Assistant

Claude reads runs, races and the backtest, searches the web for polls and news, and proposes actions. Nothing changes until you approve it here; the server then runs it through the same validated code as the Races tab. The assistant never sets a model parameter or computes a forecast number.

Checking configuration...

Synthetic Surveys

Design an instrument, field it to a probability sample of synthetic registered people with a simulated field process, and export the respondent file. Synthetic respondents only. Valid for instrument testing, weighting rehearsals and fidelity comparisons against real polls; never a measurement of public opinion. See docs/synthetic_surveys.md.

1. Instrument

Blueprint JSON (edit or paste, then "Use pasted JSON")

2. Sampling plan, field simulation and responder

Stratified by age x education x sex x race, proportional allocation.

Studies

Ask the Explorer

Guided exploration. No model integration is configured.

Choose a suggested action to update the current view.