Data, Research, and Software
Statistics and Research Methods
Prediction Stability
Queue Shift: Updating Classifiers Under a Staffing Budget
Replication Materials
When classifier predictions route cases to specialized queues, a model update changes both accuracy and the work each team receives. Queue Shift chooses the batch assignment that maximizes candidate-model probability subject to budgets on work moved between queues, counted in cases, hours, or money, and on expected negative flips. At the movement and expected flips of any comparison rule, it has at least that rule's expected accuracy under the candidate's probabilities. Simulations show that holding queue totals fixed costs little when an update mostly reorders cases and a great deal when it moves volume between queues, and that learned queue prices nearly match exact assignment in accuracy but miss the movement budget in most batches.
Stable CART: Lower Bootstrap Prediction Variance CART
Standard CART decision trees are unstable—small changes in training data can produce substantially different tree structures and predictions. Stable CART addresses this by trading a small amount of accuracy for lower cross-bootstrap prediction variance through techniques like honest estimation (using separate data subsets for learning tree structure vs. estimating leaf values), lookahead search (considering multiple future splits before committing rather than making greedy single-step decisions), and bootstrap-aware split selection (penalizing or filtering out splits that are unstable across resampled datasets).
Bagged FSR: Rehabilitating Forward Stepwise Regression
Forward Stepwise Regression (FSR) is hardly used today. That is mostly because regularization is a better way to think about variable selection. However, part of the reason for its disuse is that FSR is a greedy optimization strategy with unstable paths. Jigger the data a little, and the search paths, variables in the final set, and the performance of the final model can all change dramatically. The same issues, however, affect another greedy optimization strategy—CART. The insight that rehabilitated CART was bagging—building multiple trees using random subspaces (sometimes on randomly sampled rows) and averaging the results. What works for CART should principally also work for FSR. If you are using FSR for prediction, you can build multiple FSR models using random subspaces and random samples of rows and then average the results. If you are using it for variable selection, you can pick variables with the highest batting average (n_selected/n_tried). (LASSO will beat it on speed, but there is little reason to expect that it will beat it on results.)
Bootstrap-Consistency Regularization: Training Neural Networks for Prediction Stability
Replication Materials
Neural networks trained on the same data with different random seeds or slightly perturbed training sets can produce substantially different predictions for individual examples, even when aggregate accuracy is similar. We add a consistency penalty to the training loss that penalizes prediction disagreement across bootstrap resamples of the training data, encouraging the network to find solutions whose individual-level predictions are stable under resampling.
Selecting for Stability: Choose the Model Closest to the Ensemble
Ensembles reduce prediction variance but are expensive to deploy. Rather than averaging all models at inference time, select the single model whose predictions are closest to the ensemble average. This gives you most of the ensemble's stability benefit at the cost of a single model, with a principled selection criterion that avoids arbitrary choices among models with similar aggregate performance.
Calibration and Weighting
Calibre: Probability Calibration and Evaluation in Python
calibre: Probability Calibration and EvaluationRelated: Choosing Calibration for Decisions
Calibre combines probability calibration methods with tools for evaluating accuracy, reliability, and score ordering. A reproducible benchmark compares isotonic and centered isotonic regression, monotone splines, and other methods, illustrating why prediction granularity, calibration error, and predictive accuracy should be reported separately. The methods do not guarantee improved predictions or decision utility in every application.
fairlex: leximin calibration
Standard calibration minimizes average calibration error, which can hide large errors for minority groups. Leximin calibration instead minimizes the worst-off group's calibration error first, then the second-worst, and so on—applying the Rawlsian leximin criterion to the distribution of calibration quality across groups.
Rank-Preserving Calibration for Multiclass Classification
Rank Preserving Calibration of Multiclass Probabilities
Multiclass calibration methods can reorder predicted class probabilities, so the class a calibrated model ranks first may differ from the class the original model ranked first. This is problematic when downstream decisions depend on the ranking, not just the probabilities. We develop calibration methods that guarantee the within-example class ranking is preserved while still improving probability reliability.
First-Order Entropy Balancing via Dual Gradient Descent
Replication Materials
Entropy balancing finds survey weights that satisfy exact moment constraints while staying close to uniform weights in KL divergence. Standard implementations solve this via Newton's method on the dual, which requires computing and inverting a Hessian at each step. We show that first-order methods—multiplicative weight updates and dual gradient descent—converge reliably, scale to high-dimensional constraint sets, and avoid the numerical instabilities that plague second-order solvers when constraints are near-collinear.
Streaming Survey Raking With MWU and SGD
Python Package
onlinerake adjusts survey weights incrementally as observations arrive, using multiplicative weight updates (MWU) or stochastic gradient descent (SGD) to bring weighted sample margins toward population targets. It supports streaming survey weighting without repeatedly fitting a batch raking procedure.
streamcal: Streaming Probability Calibration
streamcal updates probability calibration as predictions and observed outcomes arrive. It uses multiplicative weight updates to adjust predicted probabilities toward observed outcome rates, allowing calibration to adapt as the data distribution changes.
From Scores to Signs: Pairwise Win-Rate Estimation with Calibrated LLM Judges
LLM-as-judge systems produce numerical scores, but downstream decisions often require pairwise comparisons: which response is better? Converting scores to win rates requires calibration—the mapping from score differences to win probabilities. We develop calibrated estimators for pairwise win rates from cardinal LLM judge scores, accounting for judge miscalibration and non-transitivity in preferences.
Record Linkage
setjoin: Record Linkage That Preserves Group Structure
setjoin implements group-constrained record matching: it assigns source groups to target groups, then matches records within the assigned groups. This is useful when both files have trustworthy group membership and the linkage should preserve it. Its current performance example is synthetic; gains on observed panel data have not been established.
preclink: High-Precision Record Linkage
A 7-step record linkage pipeline—preprocess, deduplicate, block, score, filter, decide, inspect—with multi-pass support for progressively relaxed thresholds. Implements Hungarian, greedy, and row-sequential decision rules over pairwise string similarity scores.
BloomJoin: Bloom Filter Based Joins
An R package implementing Bloom filter-based joins for improved performance with large datasets. Bloom filters provide a probabilistic test for set membership that can dramatically reduce the number of expensive exact comparisons needed during a join.
Data Collection
The Micro-Task Market for "Lemons": Collecting Data on Amazon's Mechanical Turk
With Doug Ahler and Carrie Roush.
Political Science Research and Methods, 2021.
Replication Materials
While Amazon's Mechanical Turk (MTurk) has reduced the cost of collecting original data, in 2018, researchers noted the potential existence of a large number of bad actors on the platform. To evaluate data quality on MTurk, we fielded three surveys between 2018 and 2020. While find no evidence of a "bot epidemic," significant portions of the data—between 25%-35%—are of dubious quality. While the number of IP addresses that completed the survey multiple times or circumvented location requirements fell almost 50% over time, suspicious IP addresses are more prevalent on MTurk than on other platforms. Furthermore, many respondents appear to respond humorously or insincerely, and this behavior increased over 200% from 2018-2020. Importantly, these low-quality responses attenuate observed treatment effects by magnitudes ranging from approximately 10-30%.
Optimal Data Collection When Strata and Strata Variances Are Known
With Ken Cor.
When the population is divided into known strata with known variances, the optimal allocation of a fixed sample budget across strata depends on stratum size, variance, and sampling cost. We derive the optimal allocation and show how much efficiency is lost by common rules of thumb like proportional allocation.
Geo-sampling: Sampling Randomly From the Streets
With Suriyan Laohaprapanon.Related: geoinference: inference from street samples
Estimating quantities like average potholes per kilometer or pedestrian density requires randomly sampling street locations within a region. Geo-sampling addresses this by downloading street network data from OpenStreetMap for a specified administrative region, splitting each street into 0.5km segments (recording the lat/long of segment endpoints), building a database of all segments, and then drawing a random sample that can be exported as a CSV or visualized on a map for field data collection.
Allocator: Optimal Itineraries For Spatially Distributed Tasks
With Suriyan Laohaprapanon.Related: ybar: mobile field data collection
Given a set of spatially distributed tasks (e.g., field survey locations, audit sites), Allocator computes optimal itineraries that minimize total travel time or distance, assigning tasks to enumerators and sequencing visits within each assignment.
reporoulette: Randomly Sample GitHub Repositories
Randomly sample GitHub repositories, optionally filtered by language, creation date, or star count. Useful for constructing representative samples of open-source projects for empirical software engineering research.
Total Error: What Classifier Validation Establishes About Exposure Comparisons
Replication MaterialsRelated: Unbiased Regression with Costly Item Labels
Classifier accuracy weights domains equally, but exposure comparisons and regression coefficients give each domain its own signed weight. Starting from the estimand, this paper derives how reused labels and correlated errors across domains affect bias and uncertainty. It gives sharp bounds on what exact confusion counts establish about a comparison, and shows what a probability audit must retain to improve on those bounds. A browsing-data stress test illustrates the limits of count summaries and the gains from estimating weighted errors directly.
Unbiased Regression with Costly Item Labels
fewlab: fewest items to label for unbiased OLS on shares
When running OLS on per-row trait shares (e.g., fraction of items in a category), labeling every item is expensive. Random labeling wastes budget on items that barely affect the regression. fewlab identifies the items with the highest statistical leverage—those that most influence the coefficient estimates—and prioritizes them for labeling, achieving unbiased OLS with a fraction of the labeling cost.
StatQA: Extract Multimodal Stats Q/A from Tables With Provenance
StatQA is a modern Python framework for automatically extracting structured facts, statistical insights, and multimodal Q/A pairs from tabular datasets. It converts raw columns and values into clear, human-readable statements paired with rich visualizations, enabling rapid knowledge discovery, CLIP-style multimodal RAG corpus construction, and LLM training.
Causal Inference
Inferring Treatment Compliance from Delivery-Window Data
In randomized experiments with imperfect compliance, the LATE requires observing treatment receipt, and the Wald estimator requires monotonicity. When receipt is unobserved but the experiment has a pre-treatment time series and a distinct delivery window, a structural break test applied to the delivery window can classify treated units into compliers, never-takers, and defiers. This yields an inferred compliance rate with closed-form bias correction, an empirical test of monotonicity via defier detection, and a characterization of the complier subpopulation that can be projected onto the control group.
Two Regressions and a Bootstrap: Regression Calibration for ML-Generated Covariates and the Nonlinear Boundary
When ML predictions are used as regressors in downstream models, prediction error is non-classical and a growing literature proposes purpose-built corrections (GMM, prediction-powered inference, IV, joint MLE). For linear downstream models, regression calibration—replacing the ML prediction with E[X | X̂, Z] estimated on a calibration sample—already eliminates the non-classical error structure under an exogeneity condition these methods also rely on. Two OLS regressions and a two-sample bootstrap give you consistent estimates with valid confidence intervals. For nonlinear downstream models (logistic, Poisson, any GLM), Jensen's inequality breaks the argument and heavier methods are genuinely needed. The linear/nonlinear boundary is the main result.
Partial Credit: Diagnosing Proxy Covariates with Validation Swaps
When a regression uses a proxy instead of the true covariate, how much does it matter—and where? The validation-swap framework answers both questions using an internal validation sample where both the true value and the proxy are observed. Swap validated truths into the proxy design matrix row by row and watch the coefficient move. The swap path traces the coefficient as a function of the fraction of validated rows swapped in: flat means the proxy is fine, steep means it isn't. A Shapley-value decomposition (SIM) identifies which rows drive the distortion, and a portable risk score defines a proxy-safe domain on the full sample.
Smooth Operator: Optimal Filtering of Event Study Estimates
Event study designs estimate period-specific treatment effects with known standard errors from two-way fixed effects regressions. We treat the coefficient sequence as observations from a local linear trend state-space model and apply the Rauch–Tung–Striebel (Kalman) smoother, using the known heteroskedastic regression standard errors as observation noise. The smoother adapts to local precision—trusting the trend model when a period's estimate is noisy and trusting the data when it is tight—reducing level MSE by 80% and derivative MSE by 98% versus raw estimates, while also providing a derivative-based parallel trends test with correct size.
Causal Debiasing for Robust Machine Learning
Spurious associations in training data arise because models learn correlations rather than causal relations. We propose a composite loss that incorporates causal and behavioral priors: an invariance penalty for label-preserving perturbations (e.g., swapping gendered pronouns should not change predictions), a directional penalty for perturbations that should change the label in a known direction, and a falsification penalty inspired by epidemiological negative controls that discourages reliance on features with no causal relationship to the outcome. The framework unifies ideas from CheckList-style behavioral testing, causal inference, and adversarial robustness into a single training objective.
Other
One Concept at a Time: Subspace-Constrained Causal Inference for High-Dimensional Treatments
Social scientists increasingly wish to reason causally about high-dimensional treatments (texts, images, prompts) while isolating the effect of a single latent concept. We treat a pretrained model as a measurement device and use minimal edit pairs to estimate a concept-tangent subspace in activation space, with diagnostics that can explicitly reject the existence of a clean subspace for a given encoder and concept. When diagnostics are favorable, activation-steering maps constrained to the subspace generate approximate minimal edits, and concept coordinates serve as treatments in a double/debiased ML estimator. The resulting estimand is a local, representation-dependent linear effect—not a universal causal effect of the underlying human concept.
incline: Estimate Trend at a Particular Point in Time in a Noisy Time Series
Estimating the trend (derivative) at a specific point in a noisy time series is difficult because naive approaches like computing differences between consecutive observations amplify noise rather than reveal the underlying signal. Incline addresses this by first smoothing the time series using either Savitzky-Golay filters (local polynomial fitting) or smoothing splines, then estimating the first or second derivative of the smoothed function at chosen points in time. The difference between naive and smoothed estimates can be substantial: in the provided example, the correlation between them is -0.47, making the choice of method consequential for applications like detecting sudden cost increases or identifying rapidly changing patient health trajectories.
Analytic‑Hessian Bandwidth Selection
Bandwidth selection for Nadaraya-Watson kernel regression and kernel density estimation typically relies on cross-validation, which is computationally expensive. This package implements an analytic bandwidth selector based on the Hessian of the cross-validation objective, giving a closed-form approximation that avoids the grid search.
A Lightweight ALS Solver for Iterative GLS
Generalized least squares requires estimating and inverting the error covariance matrix, which is expensive when the matrix is large or unstructured. This package uses alternating least squares with a low-rank factor-analytic decomposition of the covariance, making iterative GLS feasible for problems where the full covariance is too large to invert directly.
Optimal Classification Cutoffs for F1-score, etc.
Classifiers produce continuous scores; deployment requires a threshold. The optimal threshold depends on the metric you care about—F1, balanced accuracy, or a custom cost function—and the score distribution. This script computes the exact threshold that maximizes a given stepwise metric, avoiding the grid-search approximation common in practice.
pyppur: Projection Pursuit Dimension Reduction With Reconstruction Loss
pyppur is a Python package that implements projection pursuit methods for dimensionality reduction. Unlike traditional PP objectives geared toward finding 'interesting' projections (non-gaussian), pyppur focuses on finding non-linear projections by minimizing either reconstruction loss or distance distortion.
Names and Name-Based Inference
Predicting Race and Ethnicity From Sequence of Characters in a Name
With Rajashekar Chintalapati and Suriyan Laohaprapanon.
arXiv.orgRelated: Python Package for implementing the method.Press: InfoQ; AnacondaCON presentation (Video)
This repository compares models that estimate race and ethnicity labels from name patterns, using voter-registration data and external validation datasets. It includes preprocessing, character-based and neural models, and evaluation notebooks; these are statistical estimates from source labels rather than observations of a person's identity.
Sound Names: Classify Names Based on Sequence of Sounds
We compare name classification using phonetic encodings with classification using character sequences. In the reported experiments, an LSTM trained on Metaphone encodings performs worse, suggesting that spellings contain information lost by the sound encoding.
Graphic Names: Classify Names Using Google Image Search and Clarifai
An exploratory method that combines image-search results for first names with image tags to estimate gender associations. Validation uses labeled name data, while the project notes that search results are not representative population samples and names need not be unique to one gender.
Naampy: Infer Sociodemographic Characteristics from Indian Names
With Rajashekar Chintalapati and Suriyan Laohaprapanon.Related: outkast: caste-name data; Indian electoral rolls
Naampy estimates aggregate patterns associated with Indian first names using calibrated model scores or exact source-data lookups. Results include abstention and artifact provenance; the scores describe patterns in recorded labels, not an individual's gender.
Pranaam: Predict Religion From Name
With Rajashekar Chintalapati.
Pranaam estimates how strongly a name follows patterns associated with Muslim names in its training data. It returns a calibrated score, abstention information, and artifact provenance for aggregate research, rather than establishing a person's religion.
naamkaran: a generative model for names
With Rajashekar Chintalapati.
A character-level LSTM trained on Florida voter-registration names generates synthetic name-like strings. The outputs are intended for demonstrations, testing, and exploration, rather than verified names or representative samples of a population.
parsernaam: ML-assisted name parser
With Rajashekar Chintalapati.Related: clean-names: parsing and deduplication
Parsernaam uses character-level models to distinguish first-name from last-name tokens and infer their order in multi-token strings. It helps parse combined name fields, while its limited label set cannot represent every naming convention.
instate: State and Language Composition from Indian Surnames
With Rajashekar Chintalapati and Atul Dhingra.Related: Instate: Predict the State of Residence from Last Name.
With Atul Dhingra.
Instate estimates state composition associated with Indian surnames using electoral-roll lookups and a calibrated character model. It combines state shares with census mother-tongue shares to derive aggregate language compositions; these are name-pattern estimates, not a person's location or language.
Online Privacy and Security
Exposed: Shedding Blacklight On Online Privacy
With Lucas Shen.
Replication Materials
To what extent are users surveilled on the web, by what technologies, and by whom? We answer these questions by combining passively observed, anonymized browsing data of a large, representative sample of Americans with domain-level data on tracking from Blacklight. We find that nearly all users (>99%) encounter at least one ad tracker or third-party cookie over the observation window. More invasive techniques like session recording, keylogging, and canvas fingerprinting are less widespread, but over half of the users visited a site employing at least one of these within the first 48 hours of the start of tracking. Linking trackers to their parent organizations reveals that a single organization, usually Google, can track over 50% of web activity of more than half the users. Demographic differences in exposure are modest and often attenuate when we account for browsing volume. However, disparities by age and race remain, suggesting that what users browse, not just how much, shapes their surveillance risk.
Pwned: How Often Are Americans' Online Accounts Breached?
With Ken Cor.
ACM Web Science Conference, 2019.
Replication MaterialsRelated: Bob Rudis Analyzes Exposure by Breach; I Have Been Pwned: Evidence from the Florida Voter Registration Data
News about massive data breaches is increasingly common. But what proportion of Americans are exposed in these breaches is still unknown. We combine data from a large, representative sample of American adults (n = 5,000), recruited by YouGov, with data from Have I Been Pwned to estimate the lower bound of the number of times Americans’ private information has been exposed. We find that at least 82.84% of Americans have had their private information, such as account credentials, Social Security Number, etc., exposed. On average, Americans’ private information has been exposed in at least three breaches. The better educated, the middle-aged, women, and Whites are more likely to have had their accounts breached than the complementary groups.
Piedomains: Predict the Kind of Content Hosted by a Domain
With Rajashekar Chintalapati.Related: Domain Knowledge: Predicting the Kind of Content Hosted by a Domain.
With Suriyan Laohaprapanon. Complex, Intelligent and Software Intensive Systems (CISIS), 2020.; rdomains: domain classification in R
The package infers the kind of content hosted by a domain using the domain name, the textual content, and the screenshot of the homepage. We use domain category labels from Shallalist and build our own training dataset by scraping and taking screenshots of the homepage.
Pass-Fail: Using a Password Generator to Improve Password Strength
With Rajashekar Chintalapati.
We train character-level models on leaked passwords to study predictable password choices. Comparisons with Have I Been Pwned and randomly generated strings show how learned patterns, including common starting characters, can help diagnose password weakness.
How Often is Politicians' Data Breached? Evidence from HIBP
With Lucas Shen.
Replication Materials
Using 12,384 email addresses of politicians from 59 countries, we examine exposure to breaches recorded by Have I Been Pwned. A third of the politicians have at least one recorded breach, and more than one in five have sensitive information exposed; incomplete coverage of their email addresses makes these conservative estimates.
Bad Domains: Exposure to Malicious Content Online
With Lucas Shen.
Information, Communication and Society, 2026.
Replication Materials
We combine a month of observed browsing by over a thousand Americans with malicious-domain classifications. About half visited a malicious domain; demographic differences in median exposure disappear after accounting for how much people use the internet.
Know Your IP
With Suriyan Laohaprapanon.
A Python toolkit for enriching IP addresses with registration, network, geolocation, and reputation information from multiple providers. It reconciles provider outputs while recording provenance, caching results, and handling rate limits and partial failures.
virustotal: R Client for the VirusTotal API v3
An R client for VirusTotal API v3, covering reports on files, URLs, domains, and IP addresses, as well as submissions for analysis. The package includes request pacing, retries, and structured error handling.
Software and Tools
rmcp: R MCP Community Server
A Model Context Protocol server that lets compatible assistants invoke R tools for statistical analysis. Its documented tools cover regression, econometrics, time series, statistical testing, and data exploration.
🍠 tuber: Access YouTube API via RRelated: tubern: YouTube Analytics and Reporting
An R client for searching and retrieving YouTube videos, channels, playlists, comments, and captions. It also supports selected authenticated operations, including uploads, moderation, and playlist updates.
repaper: convert photo of a form to a web based form or an editable pdf form
With Bhanu Teja.Related: image-to-text: batch OCR; recognize: OCR quality assessment
Repaper converts photographed forms into editable digital forms. The documented workflow includes generating a Google Form from an image through a Python interface or command-line tool.
indicate: transliterate indic languages to english
With Rajashekar Chintalapati.
A transliteration toolkit for converting between Indic scripts and English-script representations. It supports composable word-table, local-model, and language-model backends, with batch processing and structured output.
Lost Years: Expected Number of Years Lost
With Suriyan Laohaprapanon.
Mortality rate is puzzling to mortals. A better number is the expected number of years lost. (A yet better number would be quality-adjusted years lost.) To make it easier to calculate the expected years lost, lost_years provides a convenient way to join to the SSA actuarial data, HLD data, and WHO life table data.
Weather Data by Location and Date
A Python package and web interface for retrieving historical daily weather by US ZIP code or coordinates. It uses NOAA station data, selects nearby stations, and returns weather measures in metric or imperial units.
LayoutLens: AI-Assisted UI TestingRelated: LayoutLens GitHub Action; UI Judge Benchmark
LayoutLens captures screenshots, page structure, styles, and geometry to measure changes between interface versions. Local comparisons report element-level differences and possible layout defects; model-based explanations are optional.
Preen: Python Package Development ChecksRelated: py-canon: shared Python package standards; r-canon: shared R package standards
Preen helps Python packages adopt and maintain the py-canon development standard. Its command-line tools scaffold repositories, update shared templates, check conformance, apply fixes, and guide tag-based releases.
Offprint: A Jekyll Theme for Academic Websites
A Jekyll theme for academic personal websites, with a single text column and publications rendered from structured data. It uses self-hosted fonts and system-controlled dark mode, without requiring JavaScript.
Adjacent — Related Repositories Recommender
A GitHub Action that adds related-repository recommendations to a README. It ranks repositories using shared topics, README similarity, or a combination, with controls for exclusions and the number of recommendations.
Advertiser: Promote Your GitHub Repositories on BlueSky
A GitHub Action that selects a repository from a curated list, generates a short description, and posts it with a link to Bluesky. It supports automated or manually triggered promotion of open-source projects.
Meta Science
Software Use and Credit
Downloads Are Cheap: Validated Use as a Measure of Research Software
Code Data Stata command index Look up a packageRelated: Recode: citation reminders; pip-fund: funding software
Counts of research software use allocate credit, and the available ones are cheap to produce: a download is one fetch of a file. I propose validated use, a package loaded in replication code deposited at a journal that checks it, as a costly signal, and measure it in 13,245 deposits from 75 economics and political science journal collections, which load 4,565 R, Python and Stata packages. Among packages the corpus uses, validated use and downloads agree loosely, and least where automatic installation is commonest: the rank correlation is 0.61 for Stata, 0.53 for R and 0.36 for Python. Downloads track a package's position in the dependency graph more closely than its use. The most used software prepares exhibits: estout is in 43% of Stata deposits. Stata appears in 1.4 times as many deposits as R and is missing from studies of research software because its packages are scattered across SSC, two journals' archives and authors' sites with no index across them. I build that index, 10,147 commands in 5,581 packages, and release it with the mention-level data.
user: Estimate Python Package Use on GitHubRelated: Python metrics
This project estimates Python package use by sampling public GitHub repositories and counting import statements. It provides data and a dashboard as an alternative to download counts, which can reflect automated installation rather than adoption.
Social Proof is in the Pudding: The (Non)-Impact of Social Proof on Software Downloads
With Lucas Shen.
Journal of Online Trust and Safety, 2026.
Replication Materials
Open-source software is widely used in commercial applications. Pair that with the fact that when choosing open-source software for a new problem, developers often use social proof as a cue. These two facts raise concern that bad actors can game social proof metrics to induce the use of malign software. We study the question using two field experiments. On the largest developer platform, GitHub, we buy ‘stars’ for a random set of GitHub repositories of new Python packages and estimate their impact on package downloads. We find no discernible impact. In another field experiment, we manipulate the number of human downloads for Python packages. Again, we find little effect.
Miscitations
Significant Error: Citations to Research With Publicized Statistical Errors
With Ken Cor.
Replication Materials
Papers containing a publicized statistical mistake continued to receive citations, and recorded qualifications were rare. After a 2011 neuroscience critique, flagged papers’ median annual citations rose from 5 in 2010 to 13–17 during 2012–2015; comparison papers’ medians rose from 4 to 10–11. In the citation-context sample, 94 of 95 valid completed ratings recorded no concern. Fixed-effects comparisons leave the size of a relative citation penalty uncertain. Continued citation does not establish that publicity had no effect.
Propagation of Error: Approving Citations to Problematic Research
With Ken Cor.
Replication Materials
Using more than 3,000 retracted articles and over 74,000 citations, we study how problematic research continues to circulate. At least 31% of citations occurred a year or more after retraction, and about 91% of post-retraction citations did not acknowledge a concern with the cited article.
Get Notified When Cited Article is Retracted
A GitHub Action that checks BibTeX references against Retraction Watch data and opens an issue when it finds a potentially retracted citation. It matches by DOI or bibliographic information and can run on a schedule or after bibliography changes.
Highlight Citations to Retracted Articles
An archived web application for finding citations to retracted articles in pasted APA references. It compares parsed citations with a database of retractions and highlights possible matches for review.
AutoSum: Summarize Publications Automatically and Discover Miscitations
AutoSum collects passages surrounding citations to a publication in openly accessible papers. Collating these passages provides a view of how other scholars summarize the work and helps identify possible miscitations.
Other
A Benchmark For Benchmarks
This note proposes a framework for constructing interpretable machine-learning benchmarks. It emphasizes defining the task and target population before choosing data, and assessing label quality and whether performance reflects the intended capability.
Review of "Noise: A Flaw in Human Judgment"
With Andrew Gelman.
Chance. 2024.
A review of Noise: A Flaw in Human Judgment that examines variation in judgments made from the same information. We consider what disagreement reveals about uncertainty, skill, effort, preferences, and the institutions in which decisions are made.
By the Numbers: Toward More Precise Numerical Summaries of Results
With Andrew Guess.
The Political Methodologist. 24(1): 2016.
We compare how often empirical articles in the American Political Science Review and American Economic Review include precise numerical results in their abstracts. The project provides coded data and analysis of differences in quantitative reporting.
The Review: Production and Consumption of APSR Articles
We collect article metadata from five political science journals to examine changes in authorship, article length, and readership. The data document the rise of coauthorship and the highly uneven distribution of article and abstract views.
superdf: Persistent Metadata for R Data Frames
SuperDF extends R data frames with metadata such as version, author, and notes. The metadata is designed to persist through common data operations and input/output, keeping documentation attached to the data.
Not to Code: Evidence From Static Code Analysis of Replication Scripts
We apply static analysis to R scripts in replication archives for American Journal of Political Science articles. The project documents style issues, warnings, and errors detected by lintr and provides the scripts used for the audit.
Public Services and Infrastructure
StreetSense: Learning from Google Street View
With Suriyan Laohaprapanon and Kimberly Ortleb.
arXiv.org
Replication Materials
How good are the public services and the public infrastructure? Does their quality vary by income? These are vital questions—they shed light on how well the government is doing its job, the consequences of disparities in local funding, etc. But there is little good data on many of these questions. We fill this gap by describing a scalable method of getting data on one crucial piece of public infrastructure: roads. We assess the quality of roads and sidewalks by exploiting data from Google Street View. We randomly sample locations on major roads, query Google Street View images for those locations and code the images using Amazon’s Mechanical Turk. We apply this method to assess the quality of roads in Bangkok, Jakarta, Lagos, and Wayne County, Michigan. Jakarta’s roads have nearly four times the potholes than roads of any other city. Surprisingly, the proportion of road segments with potholes in Bangkok, Lagos, and Wayne is about the same, between .06 and .07. Using the data, we also estimate the relation between the condition of the roads and local income in Wayne, MI. We find that roads in more affluent census tracts have somewhat fewer potholes.
AutoSense combines sampled street locations from OpenStreetMap with street-level imagery and computer vision for assessing street conditions. The workflow supports collection and analysis of infrastructure imagery at scale.
Get in Line: Waiting Times at the DMV
With Noah Finberg.
We collect reported waiting times for California DMV field offices to study variation by location, day, and hour. The analysis also examines relationships between waiting times and local sociodemographic characteristics.
Public Sector Salaries: Data on US Public Employees
A collection of public-employee salary data and supporting acquisition tools. The project aims to compare compensation across places and over time, with contextual information for analyses of local public spending and pay.
Knowledge and Misinformation
You Cannot be Serious: The Impact of Accuracy Incentives on Partisan Bias
With Markus Prior and Kabir Khanna.
Quarterly Journal of Political Science. 10(4), 489–518, 2015.
Online Appendix; Replication MaterialsRelated: Partisan Gaps in Retrospection are Highly Variable; Blog PostPress: Washington Monthly; Pacific Standard; The New York Times
Two national survey experiments test whether partisan gaps in reports of economic conditions reflect differences in knowledge or how people answer. Monetary incentives and appeals for accuracy reduce the gaps, suggesting that ordinary survey responses mix factual beliefs with partisan expression.
Motivated Responding in Studies of Factual Learning
With Kabir Khanna.
Political Behavior. 40(1): 79–101, 2018.
Replication MaterialsRelated: Blog Summarizing the Paper; The Innumerate American
Three experiments use accuracy incentives to distinguish biased learning from biased reporting of learned information. Incentives increase accurate reporting of politically uncongenial results, but also increase skepticism about their credibility, suggesting that bias can shift from one judgment to another.
A Gap in Our Understanding? Reconsidering the Evidence for Partisan Knowledge Gaps
With Carrie Roush.
Quarterly Journal of Political Science. 18(1), 2023.
Replication MaterialsRelated: An Unclear Gap: Partisan Cues and Evaluations of the Same Economic InformationPress: Not Another Politics Podcast (U. Chicago)
Across 162,083 responses to 187 questions on 47 surveys, partisan knowledge gaps are smaller and less consistently aligned with partisan interests than often assumed. Vague response options inflate some gaps, suggesting a role for motivated responding rather than differences in stored knowledge alone.
The Waters of Casablanca: On Political Misinformation
With Robert C. Luskin.
Replication Materials
Misinformation is a confidently held false belief, distinct from knowledge, mere belief, and ignorance. Multiple-choice items cannot tell it from guessing; asking how definitely true or false each statement is can. In four surveys that randomly assigned the two formats, the scale finds far less misinformation about well-publicized falsehoods, and much smaller partisan gaps, than multiple-choice items do.
Mis-measuring Political Knowledge? Do People Know More—or Even Less—about Politics than Commonly Thought?
With Robert C. Luskin and Daniel Weitzel.
Replication Materials
Revisionist studies argue that surveys understate political knowledge. Randomized experiments in three online surveys, and a recount of ANES and NAES items and don't-know probes, find little hidden knowledge: photos lower correct identification of officials, menus and probes add little beyond guessing, and asking how sure people are finds far less knowledge than multiple choice.
Misinformation About Misinformation: Of Headlines and Survey Design
With Robert C. Luskin, Yul Min Park, and Joshua Blank.
Replication Materials
Media-poll misinformation items rarely offer a 'don't know' and usually ask what respondents think. In a survey experiment, offering and encouraging 'don't know' cut incorrect answers from 32% to 20%, repeating the false claim in the question changed nothing, and asking how definitely true or false each statement is cut them to 7%.
Misinformed About the Affordable Care Act? Leveraging Certainty to Assess the Prevalence of Misinformation
With Josh Pasek and Jon Krosnick.
Journal of Communication. 65(4): 660–673, 2015.
Supporting Information | Replication Materials
Two national surveys distinguish incorrect answers about the Affordable Care Act from confidently held incorrect beliefs. Most respondents are uncertain; confident misperceptions are less common and are concentrated on provisions about which elites repeatedly made false claims.
Guessing and Forgetting: A Latent Class Model for Measuring Learning
With Ken Cor.
Political Analysis. 24(2): 226–242, 2016.
Replication Materials · Maintained Replication
Review: '... a real contribution to the literature.' — Ed HaertelRelated: R Package for implementing the method.
We develop a latent class model to estimate learning when correct answers can result from guessing. Under an assumption of no knowledge loss over short periods, simulations recover learning more accurately than raw score differences; an application to Deliberative Polls raises learning estimates and reduces measured gender gaps.
Measuring Learning in Informative Processes
With Robert Luskin and Ariel Helfer.
Replication Materials
Observed knowledge gain is a poor measure of what people learn from campaigns and deliberative forums: questionnaires ask easy items, so those who learn the most have the least room to show it. When the knowledgeable learn more, post-process knowledge is guaranteed to track true learning and observed gain is not. In simulations matched to 21 Deliberative Polls, observed gain barely correlates with true learning, and the choice of measure changes who appears to learn.
Deliberation and Learning: Evidence From Deliberative Polls
With Robert C. Luskin and James S. Fishkin.
Replication Materials · Source Data
Across 27 Deliberative Polls, factual knowledge increased by an average of 14.3 percentage points, a relative gain of 39.3%. In two polls with control groups, attendees gained an average of 11.2 percentage points more than controls. Measurements before arrival, on arrival, and after deliberation in two studies show substantial gains before participants arrive. The gender gap narrows but remains. Across 20 polls with poll-specific education cutoffs, the gap between more and less educated attendees widens.
Deliberation and Misinformation
With Robert C. Luskin.
Replication Materials
Across 19 misinformation-type items in seven Deliberative Polls, incorrect answers fell from 23% to 10%, and participants who began wrong were as likely to end right as those who began not knowing. In polls with control groups, attending cut overestimates of the number of undocumented immigrants by 15 points, but reduced confident denial of human-caused warming only slightly and not durably.
A Measurement Gap? Effect of the Survey Instrument and Scoring on the Partisan Knowledge Gap
With Lucas Shen and Daniel Weitzel.
Public Opinion Quarterly. 2025.
Replication Materials
Research suggests that partisan gaps in political knowledge with partisan implications are wide and widespread. Using a series of experiments, we investigate the extent to which partisan gaps in commercial surveys are a result of differences in beliefs than motivated guessing. Knowledge items on commercial surveys often have features that encourage guessing. We find that removing such features yields scales with greater reli- ability and higher criterion validity. More substantively, partisan gaps on scales without these “inflationary” features are roughly 40% smaller. Thus, contrary to Prior, Sood and Khanna (2015), who find that the upward bias is explained by the knowledgeable deliberately marking the wrong answer (partisan cheerleading), our data suggest, in line with Bullock et al. (2015) and Graham and Yair (2023), that partisan gaps on commercial surveys are strongly upwardly biased by motivated guessing by the ignorant. Relatedly, we also find that partisans know less than what toplines of commercial polls suggest.
An Unclear Gap: Partisan Cues and Evaluations of the Same Economic Information
With Carolyn E. Roush.
Replication Materials
Two survey experiments present the same small declines in unemployment and inflation under either an Obama cue or a Republican Congress cue. Respondents assess whether conditions got better, stayed about the same, or got worse. On a 0–100 evaluation scale, opposing-party rather than own-party cues change unemployment and inflation evaluations by -9.5 points (95% CI [-12.5, -6.5]) and -6.0 points (95% CI [-9.2, -2.8]) on MTurk, and by -5.2 points (95% CI [-10.6, 0.2]) and -8.6 points (95% CI [-14.2, -3.1]) on Lucid. The Lucid unemployment estimate includes zero. These are effects on evaluations of supplied information, not direct measures of factual knowledge; response-option vagueness was not randomized.
Measuring Perceptions of Numerical Strength of Salient and Stereotypical Groups
With Doug Ahler.
Misinformation and Mass Audiences. 2018. University of Texas Press.
Appendix
This chapter examines how to measure beliefs about the numerical size of salient and stereotypical groups. It connects systematic overestimation of stereotypical traits to broader questions about group perceptions, prejudice, and the interpretation of survey estimates.
Decision Making
Group Affect
Inter-group Prejudice
This note organizes explanations for inter-group prejudice, including group identity, personality, motivated beliefs, and differences in reasoning. It distinguishes possible mechanisms behind aversive beliefs and feelings rather than treating prejudice as a single process.
Affect, Not Ideology: A Social Identity Perspective on Polarization
With Shanto Iyengar and Yphtach Lelkes.
Public Opinion Quarterly. 76(3), 405–431, 2012.
Replication MaterialsRelated: Sort of Sorted But Definitely Cold; The Order of Feelings; Affectively Polarized?; Party TimePress: The New York Times; The Washington Post; Mother Jones; Vox, etc.
We examine polarization as social distance and hostility between partisans, rather than only disagreement over policy. Evidence from several sources shows growing dislike of opposing partisans, with policy attitudes providing an inconsistent explanation and campaign messages offering a plausible alternative.
The Parties in our Heads: Misperceptions About Party Composition and Their Consequences
With Doug Ahler.
The Journal of Politics. 80(3), 964–981, 2018.
Replication MaterialsRelated: The Partisans in our Heads; Data and ScriptsPress: FiveThirtyEight; Vox; The Washington Post; The Washington Post (2); Christian Science Monitor; The Hill; PBS (Twin Cities)
Americans greatly overestimate how many party supporters belong to stereotypical groups. Survey experiments suggest that these misperceptions are not simply expressive responding or ignorance of population shares; correcting beliefs about the opposing party reduces perceived extremity and social distance.
Typecast: A Routine Mental Shortcut Causes Party Stereotyping
With Doug Ahler.
Political Behavior. 2022.
Replication Materials | AppendixPress: Heterodox Academy
We test whether the representativeness heuristic contributes to party stereotypes. Experiments involving conjunction judgments, information about group partisanship, and cognitive load suggest that this mental shortcut can distort perceptions of who belongs to each party.
All in the Eye of the Beholder: Partisan Affect and Ideological Accountability
With Shanto Iyengar.
In The Feeling, Thinking Citizen: Essays in Honor of Milton Lodge. 2018.
Replication MaterialsRelated: Still Close: Perceived Ideological Distance to Own and Main Opposing Party; 2012 Blog PostPress: The New York Times
Why do moderate partisans remain warm toward more extreme leaders of their own party? We find both distorted perceptions of leaders' positions and asymmetric treatment of policy disagreement: even when informed about positions, partisans do not necessarily penalize more distant co-partisan candidates.
Coming to Dislike Your Opponents: Candidate Evaluations during Presidential Campaigns
With Shanto Iyengar.Related: Code and data · Replication files (v1.0.0)Press: New York Times
Candidate evaluations diverge over the 2000, 2004, and 2008 presidential campaigns, with differences across elections and measures. Ratings of both one’s own candidate and the opponent contribute. Trait-rating gaps widen faster in battleground states, while overall favorability shows a less consistent geographic pattern. These survey comparisons describe change without identifying the effects of campaigning or negative advertising.
Partisan Vision? Partisan Bias in Simple Visual Evaluations
With Carrie Roush and Alex Theodoridis.
Replication Materials
Two experiments and a survey test whether partisan cues bias simple visual evaluations, such as counting errors or assessing a scene. Across these tasks, the estimated partisan effects are generally small.
The Hostile Audience: The Effect of Access to Broadband Internet on Partisan Affect
With Yphtach Lelkes and Shanto Iyengar.
American Journal of Political Science. 61(1): 5–20, 2017.
Replication MaterialsPress: The Guardian
We use variation in state right-of-way regulations to study the effect of broadband access on partisan hostility. Combining broadband availability with survey data from 2004 and 2008, we find greater hostility and greater consumption of partisan media where broadband access expands.
Holier Than Thou? No Large Partisan Gap in Consumption of Pornography Online
With Lucas Shen.
Journal of Quantitative Description. 2024.
Replication Materials
Observed browsing data show that online pornography consumption is concentrated among a small share of users. Republicans consume somewhat more than Democrats in the unadjusted data, but the difference disappears after accounting for age and gender.
Hidden Racial Prejudice? Impact of Social Desirability Pressures on Endorsement of Racial Stereotypes
With Jon Krosnick, Tobias Stark, and Floor van Maaren.
Sociological Methods and Research.51(2), 605–631, 2019.
Replication Materials
We compare White respondents' racial stereotype reports in confidential computer-assisted and oral interviews in the 2008 American National Election Study. Confidential responses are slightly more negative about both Black and White people and have no greater predictive validity, casting doubt on a large social-desirability distortion in the oral measures.
Principled Partisan Affect? Evidence from Two Survey Experiments
With Douglas J. Ahler and Carolyn E. Roush.
Replication Materials
Is partisan affect principled? In a 2013 MTurk experiment, reports of declining rather than sustained support for a leader after controversy lowered estimated warmth toward fellow partisans, suggesting a possible preference for loyalty. In a 2018 YouGov experiment, portraying opponents as fair-minded modestly improved estimated trait ratings, with little estimated movement in warmth. Portraying allies as biased produced no clear penalty, but many readers also reported that their party was about as biased as expected. The evidence is suggestive: uncertain estimates, missing warmth ratings, bundled article features, and incomplete comparisons prevent firm conclusions about effect sizes, mechanisms, or a partisan double standard.
Pareto Partisan? Relative Gains and Support for Public Policy
With Alexander G. Theodoridis.
Replication Materials
If people prefer more for both sides to less for both, everyone should select the plan giving both sides more. In a 2018 CCES highway choice, only 35.0% of Democrats and 27.1% of Republicans did so; the rest chose smaller allocations that left their side ahead. In a post-election experiment, raising opponents’ income gain from 3% to 7%, while keeping the own-party gain at 5%, changed policy support by -20.0 points on a 0–100 scale (95% CI [-23.9, -16.1]). Relative partisan gains affect support, but the exercises do not distinguish spite from dislike of falling behind; the highway comparison also changes total spending.
Deliberation
What Would Dahl Say? An Appraisal of the Democratic Credentials of the Deliberative Polls and Other Mini-publics
With Ian O'Flynn.
Deliberative Mini-Publics. 41–58, 2014. ECPR Press.
We assess deliberative polls and other mini-publics using Robert Dahl's criteria for a democratic process: inclusion, effective participation, enlightened understanding, voting equality, and control of the agenda. The chapter examines where these institutions meet the criteria and where their democratic claims require qualification.
How Can You Think That? Deliberation and the Learning of Opposing Arguments
With Robert C. Luskin and James S. Fishkin.
Replication Materials
Using open-ended responses from a Deliberative Poll on schooling in Northern Ireland, we compare the reasons articulated by returning participants and contemporaneous controls. Participants gave somewhat more reasons, but the estimate is imprecise and participation was self-selected. The samples differed little in directional balance. Among returnees, repertoires moved toward balance chiefly because supporting reasons declined, and the estimate is sensitive to coder choice.
Deliberation and Learning: Evidence From Deliberative Polls
With Robert C. Luskin and James S. Fishkin.
Replication Materials · Source Data
Across 27 Deliberative Polls, factual knowledge increased by an average of 14.3 percentage points, a relative gain of 39.3%. In two polls with control groups, attendees gained an average of 11.2 percentage points more than controls. Measurements before arrival, on arrival, and after deliberation in two studies show substantial gains before participants arrive. The gender gap narrows but remains. Across 20 polls with poll-specific education cutoffs, the gap between more and less educated attendees widens.
Deliberation and Misinformation
With Robert C. Luskin.
Replication Materials
Across 19 misinformation-type items in seven Deliberative Polls, incorrect answers fell from 23% to 10%, and participants who began wrong were as likely to end right as those who began not knowing. In polls with control groups, attending cut overestimates of the number of undocumented immigrants by 15 points, but reduced confident denial of human-caused warming only slightly and not durably.
Deliberative Distortions? Homogenization, Polarization, and Domination in Small Group Deliberations
With Robert Luskin, Kyu Hahn, and James Fishkin.
The British Journal of Political Science. 52(3), 1205–1225, 2022.
Replication Materials
Across 2,601 group-issue pairs in 21 Deliberative Polls, we test whether discussion routinely homogenizes or polarizes attitudes, or shifts them toward those of socially advantaged participants. We find no strong or routine pattern of these distortions; the modest patterns include some homogenization and moderation.
Steadier, Not Closer: Separating Convergence from Crystallization in Deliberative Polls
Proof and Simulation · Source Data
Observed within-group variance mixes true convergence with changing measurement error. Lower error at T2 can manufacture homogenization and raise the intraclass correlation, but classical independent error cannot increase the absolute variance of group means. The paper proves these boundaries and specifies a latent partially nested model that separates convergence, between-group divergence, and crystallization. The full empirical model has not yet been fitted.
What Future for Kirkuk? Evidence from a deliberative intervention
With Ian O'Flynn, Jalal Mistaffa, and Nahwi Saeed.
Democratization. 26(7), 1299–1317, 2019.
Replication Materials | Supporting Information
A survey and deliberative field experiment examine young, educated Kirkuk residents' preferences for governing their divided society. Participants support an equal say for ethnic groups, and informed deliberation broadens support for regional autonomy without substantially changing support for equal say.
Information Environment
Not News: Provision of Apolitical News in the British News Media
With Suriyan Laohaprapanon.
Replication Materials
We classify roughly 5.4 million webpages from 276 British news outlets to estimate how much coverage concerns topics outside public affairs. A classifier trained using article URL categories and checked against hand coding estimates that about 39% of articles are apolitical news.
Strength in Numbers: Multiple Measures of Media Ideology
With Philip Habel.
Replication Materials
We combine text-based and audience-based measures of British news outlets' ideology. Parliamentary speeches, party manifestos, and Twitter following patterns provide complementary information about ideological position and differences between measurement approaches.
Measuring Agendas and Positions on Agendas
With Andrew Guess.
Replication Materials
We distinguish which issues news outlets cover from the positions they take within those issues. A supervised topic model and congressional-speech-based slant measures reveal systematic differences across 255 news programs on 50 television channels.
Notnews: Predict the Type of News Based on Story Text and URL
With Suriyan Laohaprapanon.
A Python library for classifying news by topic and distinguishing hard from soft news. It offers URL-pattern rules, trained US and UK text models, and optional language-model classification through Python and command-line interfaces.
Unreadable News: How Readable is American News?
With Lucas Shen.
We compare readability across New York Times articles and CNN, NPR, and MSNBC transcripts. The analysis examines variation across outlets and over time as one possible barrier to accessing political information.
Follow Your Ideology: A Measure of Ideological Location of Media Sources
With Pablo Barberá.
We estimate the ideological positions of more than 2,300 media sources using the following patterns of politically interested social-media users. Comparisons with content-based measures support the approach, and applications examine journalists' ideology and ideological diversity within outlets.
The Supply of Media Slant Across Outlets and Demand for Slant Within Outlets: Evidence from US Presidential Campaign News
With Marcel Garz, Daniel Stone, and Justin Wallace.
European Journal of Political Economy.
Replication Materials
We study presidential campaign horse-race headlines across six online news outlets in 2012 and 2016. Headlines tend to favor their outlets' typical readers, but within-outlet engagement provides little evidence that readers prefer congenial headlines and somewhat more evidence of demand for uncongenial ones.
Don't Expose Yourself: Discretionary Exposure to Political Information
With Yphtach Lelkes.
Oxford Research Encyclopedia of Politics. 2018.
Replication MaterialsRelated: Categorizing the Content of Domains; Measuring Selective Exposure; The Fairest of All
We reconsider claims that greater choice over political information increases knowledge gaps and polarization. The review questions the premise of widening knowledge gaps and finds that the evidence connecting selective exposure to polarization is more limited and nuanced than common accounts suggest.
The Good NYT: Provision of Apolitical News in the New York Times
We use the annotated New York Times corpus from 1987 to 2007 to study the mix of public-affairs and apolitical coverage. The project also develops analyses of geographic focus using article metadata and location fields.
Hard News: The Softening of Network Television News
With Daniel Weitzel.
Replication Materials
A coded sample of more than 5,000 network television news segments from 1968 to 2019 documents changes in political content and geographic focus. Coverage unrelated to politics increases over time, and local news becomes a substantially larger share in the later years.
Partisan Imbalance in Politifact?
We combine scraped PolitiFact ratings with independently coded party affiliations to examine differences in scrutiny and ratings across parties. The analysis distinguishes how often statements are checked from the average rating among the statements selected for checking.
Top News! URLs from News Feeds of Major National News Sites (2022-)
With Derek Willis.
A scheduled collector archives URLs from major news sources for studying news production. The repository contains source-specific URL collections and links to historical full-text releases, whose coverage differs from the current collections.
CNN Transcripts 2000--2025Related: Internet Archive TV captions; Vanderbilt TV news abstracts; NBC / MSNBC transcripts
A corpus and scraper for CNN transcripts spanning 2000 through March 2025, with program names, dates, times, and transcript text. The repository documents collection across versions of CNN's site and links to research-access data on Harvard Dataverse.
What and Who is on Network Television?
We combine television schedules with program metadata to study changes in genres and the race and gender composition of casts and production teams. The repository includes collection scripts, data, and figures describing these trends.
Working Women on Indian TV
With Asha Sood.
We code employment among fictional female characters in Indian evening television soaps. In data collected in 2015 on 172 characters across 20 shows, 36.6% of working-age female characters are portrayed as working, with substantial variation across shows.
The Face of Crime in Prime Time: Evidence from Law and Order
With Daniel Trielli.
Replication MaterialsPress: The Washington Post
We code the race and gender of victims and criminals in three Law and Order series and compare portrayals with crime statistics. White people and women are overrepresented in both roles, while Black people and men are underrepresented.
Other
Problem Solving
A practical framework for diagnosing problems: generate specific explanations, assess their plausibility, and design checks that distinguish among them. Examples show how correlations, close examination of failures, and comparisons with successful cases can help identify causes.
Is an Uncertain Prospect Less Preferred Than Its Worst Possible Outcome? New Evidence on the Uncertainty Effect
With Doug Ahler.
Replication Materials
Two larger surveys replicate the finding that people sometimes prefer a certain option to a lottery whose worst outcome is better. Clarifying the guarantee reduces the effect, while expressing the lottery as natural frequencies or using monetary choices eliminates it in these experiments.
Description Invariance in Risky-Choice Framing: Evidence from Five Survey Experiments
With John Protzko and Jon Krosnick.
Replication Materials
Five archival survey experiments test whether equivalent gain and loss descriptions change risky choices. All 23 contrasts show the classic framing direction, but explicitly stating the certain outcome removes much of the gap, suggesting that inferences from incomplete descriptions matter alongside gain and loss labels.
Is Cynicism Taken as Evidence of Political Sophistication?
With Alexander G. Theodoridis.
Replication Materials
In a two-arm comparison in the 2020 Cooperative Election Study, 711 respondents rated one of two tweets. Both said the system rewards bad behavior; one said politics attracts good people, the other bad people. The bad-people tweet scored 6.2 points lower in perceived sophistication on a 0–100 scale (95% CI: −9.4 to −3.0). Covariate adjustment and survey weighting preserve the negative difference. A causal interpretation requires random assignment within the rating sample; the original questionnaire and randomization program are unavailable.
Mixed Signals: Movie Quality Assessments Across Platforms
We compare movie ratings across platforms using data on 16,319 American films released between 1950 and 2020. Among platforms with sufficient coverage, average ratings are only moderately correlated, revealing substantial differences in assessments of the same films.
Americans' Attitudes Toward The Affordable Care Act: Would Better Public Understanding Increase or Decrease Favorability?
With Wendy Gross, Tobias Stark, Jon Krosnick, Josh Pasek, Trevor Tompson, Jennifer Agiesta, and Dennis Junius.Press: Forbes; Pacific Standard; The Dish, among other outlets.
Surveys from 2010 and 2012 find substantial misunderstanding of the Affordable Care Act despite support for many of its provisions. The analysis estimates that fuller understanding could increase support for the law; this is a modeled implication, rather than a demonstrated effect of an information campaign.
Americans' Attitudes toward the Affordable Care Act: What Role Do Beliefs Play?
With Gabriel Miao Li, Josh Pasek, Jon Krosnick, Tobias Stark, Jennifer Agiesta, Trevor Tompson, and Wendy Gross.
Annals of the American Academy of Political and Social Science. 2022.
We examine whether Americans evaluate the Affordable Care Act through beliefs about its individual provisions. Despite partisan disagreement, both Democrats' and Republicans' evaluations are associated with what they believe the law does and how confident they are in those beliefs.
Revisiting a Natural Experiment: Do Legislators With Daughters Vote More Liberally on Women's Issues?
With Don Green, Oliver Hyman-Metzger, and Michelle Zee.
Journal of Political Economy Microeconomics. 2023.
Replication Materials | Supporting InformationPress: Phys.org
We revisit evidence that legislators with daughters vote more liberally on women's issues, extending the analysis to eight congresses before and eight after the original study period. We find no daughters effect in either extension, challenging the explanation that later party polarization alone erased the relationship.
Extreme Recall: Which Politicians Come to Mind?
With Daniel Weitzel.
Journal of Elections, Public Opinion and Parties. 2024.
Replication MaterialsRelated: Extreme Recall
Two surveys nearly a decade apart show that people's images of political parties center on a few prominent national figures, with more extreme politicians recalled more often. Television news coverage is associated with which politicians come to mind.
Political Economy of India
The Limits of Electoral Gender Quotas in Rural Local Bodies
With Varun Karekurve-Ramachandra.
Replication MaterialsRelated: Governance; Public spending; PDS; MarriagePress: VoxDev
Across more than 34,000 rural local bodies, prior reservation produces small or inconsistent short-run gains in women's success in subsequently open seats. Phone audits document extensive male mediation of access to women elected under quotas, a pattern consistent with, but not proof of, proxy governance limiting independent political careers.
Female Reservation and the Qualifications of Local Elected Officials in India
Replication Materials
Schooling differences under women's reservation follow distinct patterns across offices and settings. Published village-head studies and our northern village-head estimates show deficits among reserved-seat winners; Kerala's village ward members and councillors in Mumbai and Delhi show no consistent deficit, although the urban estimates do not establish equivalence. We synthesize these groups separately, rather than treating a single overall average as the main finding. Reserved-seat winners are more likely to report no occupation in the village, block and city offices with occupation records, but face fewer pending criminal charges in Mumbai and Delhi. The records do not establish effects on governing performance.
The Older Half: Spousal Age Gap in India
With Suriyan Laohaprapanon.Related: Indian electoral rolls
Using electoral-roll records for nearly 70 million couples across Indian states and union territories, we describe spousal age differences. Husbands are generally older, with an average gap of about 4.1 years and substantial variation across states and spouses' ages.
Unlanded: Distribution of Land in Bihar
With Lucas Shen.
Replication Materials
Digitized land records from Bihar reveal concentrated ownership and substantial disparities by gender, religion, and caste. Men make up 76.5% of landowners, and Muslims, Scheduled Castes, and Extremely Backward Classes are underrepresented; the study measures ownership rather than operational agricultural holdings.
Indian Electoral Rolls: Data and Extraction Tools
With Atul Dhingra.Related: Searchable-roll parser; Scanned-roll parser
Scripts and linked tools collect and parse Indian electoral rolls published as PDFs by state election authorities. The project provides a foundation for research using recorded names, ages, gender labels, family relationships, and electoral geography.
Missing Women
Missing women on Indian streets
With Varun Karekurve-Ramachandra.
Replication MaterialsPress: Marginal Revolution; Watching the Wheels (podcast)
Randomly sampled streets in four Indian cities reveal a large gender imbalance in public space: women account for 12% to 16% of observed adults. Comparisons with residential populations and time-use data suggest that much of the gap arises because women leave home less often and make fewer trips.
Son Bias in the US: Evidence from Business Names
With Walter Guillioli.
Replication Materials
We compare businesses named with 'son' or 'sons' and 'daughter' or 'daughters' using state business registries. Across 40 states, the median ratio is 12 to 1, with limitations in registry searches making the analysis a conservative accounting of the imbalance.
Which Women Are Missing? Adult Sex Ratio By Last Name
With Suriyan Laohaprapanon.
We use Indian electoral-roll records for more than 35 million people to estimate adult sex ratios by surname. The analysis describes variation that state-level averages can obscure, using recorded electoral-roll gender labels.
Missing Women on the Streets
An exploratory project on measuring women's presence in public space using observations of people on streets. It develops sampling and data-collection approaches and discusses the assumptions needed to compare observed street populations with a gender-balance baseline.
Epic Children: Sex Ratio of Children of Key Characters in Epics
A dataset records sons and daughters attributed to key figures across epic and mythological traditions. It documents large differences in the visibility of named sons and daughters, with source and family-structure information for interpreting the counts.
Missing Daughters of Indian Politicians
We use official biographies of members of India's 12th through 17th Lok Sabhas to examine the sex ratio of politicians' children. The project compares recorded sons and daughters to assess whether the imbalance observed in the wider population also appears among political elites.
Cricket
Elo Ratings of International Cricket Teams By Format
With Derek Willis.Press: The Hindu
We construct historical monthly Elo ratings for international cricket teams, separately for Tests, one-day internationals, and T20 internationals. The repository provides match outcomes, rating scripts, and downloadable series for comparing team strength over time.
WAR Ratings for Cricketers
With Derek Willis.
We develop batting and bowling Wins Above Replacement measures for men's and women's one-day international cricket. The project compares simple aggregate measures with context-adjusted approaches, with results depending on the chosen replacement baseline and conversion from runs to wins.
Fairly Random: The Effect of Winning the Toss on Winning the Match
With Apoorva Lal, Derek Willis, and Avidit Acharya.
Journal of Sports Analytics. 2023.
Replication MaterialsRelated: Fairly Random: Impact of Winning the Toss on the Probability of Winning (with Derek Willis; arXiv.org).; python-espncricinfo; Cricket data and collection scriptsPress: ESPN: How much does the toss really matter?; ESPN: Why replacing the toss with an auction is the fair thing to do
We use cricket's coin toss to estimate the advantage of choosing whether to bat or bowl first. The estimated effect on winning is negligible in daytime international limited-overs matches, but about 3.3 percentage points in day-night matches and 4.9 points in Tests; rain-adjusted matches show no statistically distinguishable effect.