Skip to content

A personal data project · collected 2025 · revived 2026

What US campaigns put on the record, read closely.

In August 2025 I wrote two small Python scrapers for The American Presidency Project at UC Santa Barbara. One collected the archive's campaign documents from 2016 to 2024; the other collected every presidential and vice-presidential debate transcript from 1960 to 2024, plus the primary debates it holds.

Campaign Text Lab turns those CSV files into a reading room. It counts, compares and charts; it does not judge, predict or score anyone. The texts themselves stay on the archive, one click away from every number here.

Campaign documents

7,556

54 speakers, 2016 to 2024

Words in candidates' own voice

4.5M

after removing other speakers

Debate transcripts

179

49 general-election and VP, 130 primary

Speaking turns

44,255

2.9M words, 1960 to 2024

Campaign documents per month

1 January 2016 to 6 November 2024

The archive fills up in election years: primaries in the first half, conventions and the general election after.

0200400600800201620172018201920202021202220232024

Six ways in

Choose a question

What the 2025 scrapers produced

The original results, as collected

The summary the 2025 notebook printed, from its own fields and all of its listing rows. The rest of the site counts each document once.

Documents by type

Assigned from title words by the scraper. All 7,582 listing rows, as collected; 26 of them list a page a second time, so the site works with 7,556 documents.

  • Document5,4965,475 once each
  • Statement998
  • Speech/Remarks554552 once each
  • Debate411409 once each
  • Address6564 once each
  • Interview58

Documents by candidate

The archive's own filing, as collected; volumes reflect how much each campaign released. Top eight of 54.

  • Donald J. Trump (1st Term)1,2261,224 once each
  • Joseph R. Biden, Jr.1,0081,003 once each
  • Nikki Haley842
  • Bernie Sanders575
  • Hillary Clinton474473 once each
  • Ron DeSantis448
  • John Kasich383382 once each
  • Marco Rubio380377 once each

Ground rules

Descriptive by design

Symmetric

Every candidate is treated the same way. Colours mark groups being compared, never parties; the defaults compare election cycles, not people.

Counts, not verdicts

Word frequencies, turn lengths and reading grades describe texts. They are not measures of honesty, quality or sentiment, and nothing here forecasts anything.

The texts stay at the source

The database holds counts and metadata. Quotations are capped at 25 words and link to the archive page they come from.

Uncertainty on show

Rates, grades and agreement scores come with 95% intervals, sample sizes and the seed that reproduces them, and word lists come with a stability check.

AI is optional and logged

The one AI feature runs only with your own key, sends only short excerpts, labels its output and records every call in an audit log in your browser.

Decisions written down

Each design choice has a decision record: what was decided, the options, what happened and what I would change.

About this project

A personal project, 2025

Campaign Text Lab started as a personal data-collection exercise in August 2025: two Jupyter notebooks that page through the archive politely (rate-limited, with retries, a pickle cache and a resumable checkpoint), parse each page with BeautifulSoup and write tidy CSVs. It was not coursework and is not affiliated with The American Presidency Project.

In 2026 it was revived as this website. Nothing was re-scraped: a reproducible script reads the original CSVs and writes a small read-only SQLite database of derived statistics, which the site queries on the server. The original parsing code was ported to TypeScript and is tested against the original outputs, row for row.

Provenance. The notebooks and CSVs are preserved unchanged in the repository's original/ folder. Source texts: Gerhard Peters and John T. Woolley, The American Presidency Project, University of California, Santa Barbara. Copyright © The American Presidency Project.

Type
Personal project · 2025 (revived 2026)
Author
Sunchuangyu (Rin) Huang
Original stack
Python, Jupyter, requests, BeautifulSoup4, pandas, ThreadPoolExecutor, tqdm
Revived stack
Next.js 16, React 19, TypeScript, Tailwind CSS 4, SQLite via node:sqlite, Python build scripts run with uv, Vitest