Skip to main content
Political Science· 15-page report· 1 figure

2020 County Vote-Share Demographic Regression

Regress 2020 county vote share on demographics and validate out-of-sample predicted versus actual results.

What this research found

How much of the 2020 presidential vote is written into county demographics? Official county returns were merged with American Community Survey five-year estimates for 3,111 counties, and the Republican share of the two-party vote was regressed on just five covariates: college attainment, non-Hispanic white share, median household income, population density, and median age. The model explains 64.6% of the cross-county variance, and when it is retrained without ever seeing ten of the fifty states it still explains 45% of the variance in those states.

  • Five demographic covariates account for 64.6% of the cross-county variance in Republican two-party share (R² = 0.6455, adjusted 0.6449, F = 1,130.8) with an in-sample root-mean-squared error of 0.097 across 3,111 counties.
  • Education and racial composition dominate, and almost cancel each other out. One standard deviation more bachelor's-degree holders (about 9.7 percentage points) is associated with an 8.5-point lower Republican share, while one standard deviation more non-Hispanic white residents (about 20 points) is associated with an 8.4-point higher share.
  • Population density adds an independent negative pull of −0.052 per standard deviation even with education, race, income, and age held fixed. Median age (−0.025) and log income (+0.027) are the weakest contributors.
  • Clustering standard errors by state inflates them by factors of 2.3 to 6.1 relative to classical ones — most for the geographically clustered predictors, non-Hispanic white share (6.13 times) and log density (5.00 times). All five predictors remain significant at p < 0.01 under the conservative errors.
  • Trained on 2,433 counties in 40 states and tested on 678 counties in the 10 withheld states, the model reaches a held-out R² of 0.4494 and an RMSE of 0.1022 — predictions typically within about ten percentage points. Training-set R² of 0.676 sits close to the full-sample value, so the gap reflects extrapolation across unseen state political cultures rather than overfitting.
  • The largest errors expose the cost of one racial covariate. Miami-Dade County came in at 0.46 Republican share against a predicted 0.21, while Native American-majority counties went the other way: Sioux County, North Dakota at 0.24 actual versus 0.59 predicted, and Menominee County, Wisconsin at 0.18 versus 0.52.

How it was done

County returns for 2020, ACS 2016–2020 five-year demographic estimates, and the 2020 Census Gazetteer land-area file were joined on five-digit FIPS codes, with identifiers zero-padded to preserve leading zeros, retired and renamed codes remapped, and non-participating territories removed. The join retained 98.7% of participating election jurisdictions; the systematic non-matches are Alaska's forty state-legislative districts, which report presidential results without a county-level demographic analogue. Income and population density were log-transformed to tame right skew spanning nearly seven-fold and five orders of magnitude respectively, and all five predictors were z-scored so coefficients read as effect sizes. Ordinary least squares was fitted with state-level cluster-robust standard errors, and generalization was tested by shuffling the fifty states under a fixed seed into a 40-state training set and a 10-state test set, with standardization statistics taken only from the training counties.

Data sources

  • MIT Election Data and Science Lab — County Presidential Election Returns 2000–2020 (3,152 county-equivalent jurisdictions in the 2020 contest)
  • US Census Bureau American Community Survey 2016–2020 five-year estimates — nine raw variables at the county summary level
  • US Census Bureau 2020 Gazetteer file — county land area for the density calculation

Limitations

This is an ecological analysis: the coefficients describe associations among county aggregates and cannot be read as individual-level effects. A single non-Hispanic white share collapses politically distinct Black, Hispanic, and Native American populations into one complement, residuals remain mildly heteroskedastic and heavy-tailed, and ACS margins of error are treated as zero.

How this research was produced

K-Dense Web planned and ran this political science investigation end to end — gathering the sources, carrying out the analysis, producing the figures, and drafting the report. The full session transcript, including every intermediate step, is available to view.

Share:
Social Science

East vs West Cultural Analysis

Quantify individualism-collectivism cultural differences using Hofstede dimensions and World Happiness Report data.

Demographics

Indian State Emigration Patterns

Analyze state-wise emigration drivers with diaspora networks explaining 78% of variance using panel regression.

Scientometrics

India vs China AI Research Analysis

Compare AI research productivity and impact between India and China using OpenAlex bibliometric data (2019-2024).

Run this kind of analysis on your own question

Try K-Dense Web free and see how an AI co-scientist accelerates your research.