Skip to content
AI.info

Research

The Human Flourishing Geographic Index: A County-Level Dataset for the United States, 2013--2023

Overview Research area: Computational social science and natural language processing, specifically the use of large language models to measure subjective well-being from social media text. The paper i

The Human Flourishing Geographic Index: A County-Level Dataset for the United States, 2013--2023
arXiv
2511.03915
Published
2025-11-05
Authors
Stefano M. Iacus, Devika Jain, Andrea Nasuto, Giuseppe Porro, Marcello Carammia, Andrea Vezzulli

AI summary

Overview

  • Research area: Computational social science and natural language processing, specifically the use of large language models to measure subjective well-being from social media text. The paper is primarily a dataset/measurement paper rather than a modeling-architecture paper.
  • Technical level: Intermediate. The classification pipeline, the aggregation formulas, and the regression validation are described in enough detail to follow, but the work assumes familiarity with LLM fine-tuning, geospatial joins, and basic regression.
  • Scope in one sentence: The paper introduces the Human Flourishing Geographic Index (HFGI), a county- and state-level, monthly and yearly dataset of flourishing-related indicators derived from roughly 2.6 billion geolocated U.S. tweets covering 2013–2023.

What This Paper Is About

Measures of human well-being such as GDP or national happiness reports are usually too coarse in space and time to show how flourishing varies between neighboring counties or shifts month to month. The authors aim to fill that gap by classifying the text of about 2.6 billion geolocated U.S. tweets with fine-tuned large language models and aggregating the results into 48 flourishing-related indicators at county and state level, at monthly and yearly frequency. The result is a public dataset intended to support research on well-being, inequality, and social change at much finer resolution than surveys allow.

Key Contributions

  1. A new fine-grained well-being dataset. The HFGI publishes 48 indicators — 46 drawn from Harvard's Global Flourishing Study (GFS) questionnaire framework plus measures of attitudes towards migration (migmood) and perception of corruption — at county and state level, at monthly and yearly resolution, with tweet counts, valid tweet counts, salience, and standard deviations for each cell.
  2. A scalable LLM labeling pipeline. Tweets are labeled by a fine-tuned Llama 3.2 3B model that returns a compact JSON dictionary naming only the well-being dimensions present in a tweet and rating each as low, medium, or high. The paper argues this small-model choice trades a controlled amount of accuracy for the throughput needed to process 2.6 billion tweets.
  3. A human-supervised fine-tuning corpus. Starting from 10,000 randomly selected tweets, a team of 12 human coders produced a final corpus of 4,581 coded tweets, split into 2,404 training and 2,177 test tweets, which expands to 110,584 and 100,142 coding instances respectively because each tweet can carry all 46 dimensions.
  4. External validation against offline and administrative series. The authors compare their online believegod indicator with 2020 Religious Congregations and Membership Study (RCMS) adherence data from ARDA, and estimate rural–urban differences in expressed well-being using USDA Rural–Urban Continuum Codes.

Main Findings

  • Highest and lowest salience dimensions: Indicators tied to subjective well-being and emotion (happiness, optimism, empathy, depression) show the highest salience across 2013–2023, while civic, moral, or transcendent themes (forgiveness, volunteering, charity, afterlife) appear much less frequently.
  • Persistent geographic structure: County-level salience maps show coherent regional patterns even for low-base topics such as corruption, with contiguous "corridors" of higher discourse intensity rather than random noise. The happiness panel shows broad nationwide coverage with bands of higher salience across large parts of the interior West and the central corridor, and relatively lower salience in some metropolitan belts and parts of the coastal Northeast. belonging is moderate-to-high across the Midwest and interior, lower around several major metropolitan areas; finworry is most elevated across the Great Plains and portions of the rural Midwest.
  • Religion: online discourse tracks evangelical adherence more than general adherence. At county level (N = 3,099), believegod correlates with evangelical adherence at Pearson r = 0.385 and Spearman ρ = 0.507, but only r = 0.099 and ρ = 0.158 with total religious adherence; the two institutional measures correlate at r = 0.596 and ρ = 0.385. At state level (N = 51), online vs evangelical rises to r = 0.760 and ρ = 0.660, online vs all is r = 0.351 and ρ = 0.318, and evangelical vs all is r = 0.457 and ρ = 0.313.
  • Rural–urban contrasts: In county-level regressions on 46 flourishing indicators, indicators of religious faith and moral virtue (believegod, relcomfort, forgive, lovedgod) and of happiness and purpose (lifesat, purpose, balance, volunteer) are higher in rural counties, while urban counties show stronger expression of polvoice, delayed gratification, charity, and relcrit. Only statistically significant contrasts (unadjusted p < 0.05) are shown. The authors describe these as correlational, not causal.
  • Reported computation cost: Classifying the 2.6 billion tweets used quantized fine-tuned Llama models on mainly NVIDIA A100/H100 GPUs for a total of 1,045,048 hours, largely on spare GPU cycles across FAS-RC, Kempner Institute, and NSF-ACCESS facilities.
  • Validation results beyond the coding-corpus description are not reported in the available text. The paper states that "we observed that non–fine-tuned model" and then the provided content ends, so the comparison between fine-tuned and non-fine-tuned performance is not available here.

Methodology in Plain English

The authors start from the Harvard CGA Geotweet Archive v2.0, a collection of about 10 billion geolocated tweets worldwide gathered from January 2010 to July 12, 2023, when Academic API access ended. Only tweets carrying explicit spatial metadata were retained, roughly 1–2% of all tweets on average. From this archive they isolate 2.6 billion tweets from the United States and attach census tract variables using the U.S. Census Bureau's 2020 TIGER/Line Shapefiles.

They then take 46 questions from the Global Flourishing Study questionnaire and rewrite them as short prompts for a language model. A fine-tuned Llama 3.2 3B model reads each tweet and returns a JSON dictionary listing only the dimensions that apply, each rated low, medium, or high. Those labels become numbers: −1 for low, +0.5 for medium, +1 for high, and 0 when the dimension is absent. Two extra tasks use their own prompts: migration attitude is sorted into pro-immigration, anti-immigration, neutral, or unrelated (mapped to +1, −1, 0, and dropped), and corruption perception uses a six-question prompt where the indicator itself is built solely on the first question, with answers to the other questions released as binary variables.

Because every tweet is forced to carry a value for all dimensions, the coded data can be summed efficiently with vectorized operations in DuckDB. Sums are first taken per census tract per day, then rolled up to county and state and to monthly or yearly totals. Dividing each sum by the number of tweets in which the dimension was actually present — zeros are excluded from the denominator — gives indicators that range from −1 to +1, representing average polarity or intensity. Standard deviations of those averages are computed alongside them, and a salience measure is defined as valid tweet count divided by total tweet count for a cell.

For validation, tweets were first machine-labeled by a variety of models from small Llama versions up to ChatGPT-4. Human coders were then shown each tweet alongside all the model categories and ratings, and were asked to reject classifications and ratings that did not apply. The resulting corpus was split into training and test sets. The paper also compares the online religion indicator with county-level RCMS 2020 adherence data, and estimates rural–urban differences with county-level regressions controlling for log tweet volume and log population, using USDA Rural–Urban Continuum Codes 1–3 as metropolitan and 4–9 as nonmetropolitan.

Why This Matters

Impact on research. Survey-based well-being measurement is limited by declining response rates, observer effects, and cost, which restrict how often and how finely well-being can be measured. The HFGI offers monthly and yearly observations for essentially every U.S. county over a decade, letting researchers examine spatiotemporal dynamics, regional disparities, and local factors that foster or hinder flourishing, and to join these layers with socioeconomic, health, demographic, and CDC datasets. The authors are careful to frame the indicators as a measure of expression propensity — how often people publicly talk about a dimension — rather than population prevalence of private attitudes.

Real-world applications:

  • Mapping where discussion of financial worry, belonging, or civic concern concentrates, which can help target outreach, service provision, or further data collection to specific counties.
  • Using salience layers as "opportunity" or "visibility" weights in spatial models that compare indicator levels across places with very different information ecosystems.
  • Benchmarking local religiosity, social anchoring, or economic anxiety against institutional records such as RCMS adherence and county-level economic data.
  • Monitoring rural–urban divides in expressed purpose, faith, and civic engagement to inform regional policy and community programs.

Industry relevance. The pipeline demonstrates a practical pattern for turning enormous text corpora into interpretable, spatially and temporally indexed measurements: a small fine-tuned open-source model for throughput, structured JSON output to cut inference cost, and an embedded analytical database for aggregation at scale. The same architecture — fine-tune on a modest human-coded set, classify at scale, aggregate to fixed geographies and time windows — transfers to market research, brand and public-opinion tracking, insurance and public-health analytics, and any domain where survey data is too slow or too expensive to collect continuously.

Future Directions

  • Extending beyond the United States. The authors chose Llama 3.2 for its official support of eight languages (English, German, French, Italian, Portuguese, Hindi, Spanish, and Thai) explicitly to scale the approach to the rest of the global geotweet archive, and report that fine-tuning helps the model handle languages beyond those eight with few or no examples.
  • Completing the validation evidence. The paper's validation section, as available here, stops before reporting how the fine-tuned model compares with non-fine-tuned models on the test set, and how well the measures correlate with sentiment series, CDC mental-health indicators, and CPI. Those comparisons are needed to substantiate the claim that the indicators accurately represent the underlying constructs.
  • Testing whether low-salience spatial clusters are real. The authors state that the localized clusters in the corruption panel suggest but do not definitively establish geographic patterning, and that further analysis is needed to determine whether these reflect coherent regional structures or stochastic variation.
  • Descending below the county level. The dataset deliberately publishes only county and state aggregates at monthly and yearly frequency to protect privacy and ensure sufficient tweet volumes, leaving open whether tract-level or daily indicators could be released under stronger privacy protections.

Target Audience

Researchers in computational social science, well-being and happiness economics, public health, and geography who need spatially and temporally resolved measures of flourishing. It should also interest NLP practitioners studying how far small fine-tuned models can be pushed on classification tasks at billion-document scale, and policy analysts or data journalists who want a ready-made, openly published county-level indicator set to combine with their own data.

Authors’ abstract

Quantifying human flourishing, a multidimensional construct including happiness, health, purpose, virtue, relationships, and financial stability, is critical for understanding societal well-being beyond economic indicators. Existing measures often lack fine spatial and temporal resolution. Here we introduce the Human Flourishing Geographic Index (HFGI), derived from analyzing approximately 2.6 billion geolocated U.S. tweets (2013-2023) using fine-tuned large language models to classify expressions across 48 indicators aligned with Harvard's Global Flourishing Study framework plus attitudes towards migration and perception of corruption. The dataset offers monthly and yearly county- and state-level indicators of flourishing-related discourse, validated to confirm that the measures accurately represent the underlying constructs and show expected correlations with established indicators. This resource enables multidisciplinary analyses of well-being, inequality, and social change at unprecedented resolution, offering insights into the dynamics of human flourishing as reflected in social media discourse across the United States over the past decade.

Read the original paper