Skip to content

Method and source

How the archive was collected, checked and analysed

Everything on this site comes from the CSV files the 2025 notebooks wrote. Nothing was scraped again. This page describes the original collector, the parts of it that were ported to TypeScript, every calculation and interval the site adds, the topic-label evaluation, the assumptions and limits, the decisions behind them, and how AI is and is not used.

On this page

01

The collection pipeline (August 2025)

Two Jupyter notebooks read public listing pages on The American Presidency Project and then each linked page, with browser-like headers, connection pooling, retries with exponential backoff on 429 and 5xx responses, and fixed delays between requests.

documents.ipynb

  1. Phase 1 · listing

    CampaignDocumentsScraper

    Requests the campaign-documents category at 1,000 items per page, finds the page count (pagination links, else a binary search over page numbers), and scrapes pages in a thread pool of up to five workers: date, title, link and the “Related” candidate.

  2. Phase 2 · content

    OptimizedDocumentExtractor

    Fetches each unique link in batches of 50 across eight workers, with a pickle cache and a JSON checkpoint every 500 URLs so an interrupted run resumes. It extracts title, date, speaker, byline, paragraphs, a document type and location from the title, a video flag and a word count.

debates.ipynb

  1. Phase 1 · listing

    scrape_debates_data

    Reads the debates category at 200 items per page: date (normalised from the ISO timestamp), title, link and related category. 180 rows, 179 unique transcripts.

  2. Phase 2 · transcripts

    extract_debate_info

    Fetches each transcript one second apart, keeps the full HTML, pulls the “PARTICIPANTS:” and “MODERATORS:” blocks into lists, and flattens the text with line breaks preserved.

  1. 2026 · build

    scripts/build_analytics.py

    Run with uv. Reads the three CSVs from original/ and writes a 9 MB read-only SQLite file of derived data: no running text.

  2. 2026 · serve

    Next.js on the server

    Pages query the database with Node's built-in node:sqlite. Static pages are prerendered; tools that take parameters compute on request.

  3. 2026 · test

    Vitest in CI

    The TypeScript ports and the statistics are re-checked against the original CSVs on every push.

02

Checking the port against the original outputs

Two independent checks. scripts/parity_check.py executes the notebook cells verbatim, with the network replaced by a stub that serves pages rebuilt from the stored data. The Vitest suites then run the TypeScript ports over the same CSVs.
  • Debate transcripts → Participants, Moderators, lists, plain text

    Rows
    179 transcripts
    Original Python
    identical
    TypeScript test
    debate-splitter.test.ts
  • Debate listing → ISO timestamp to “October 01, 2024”

    Rows
    180 rows
    Original Python
    identical
    TypeScript test
    dates.test.ts
  • Documents → type, location, word count, paragraph join, date

    Rows
    7,582 rows
    Original Python
    identical
    TypeScript test
    documents.test.ts, dates.test.ts
  • Documents → listing date round-trip to the Date column

    Rows
    7,582 rows
    Original Python
    identical
    TypeScript test
    dates.test.ts
  • Revival analytics → tokens, sentences, syllables, grade per document

    Rows
    7,556 documents
    Original Python
    n/a (new in 2026)
    TypeScript test
    textkit.test.ts (TypeScript vs the Python build)
  • Fightin' Words z-scores and Poisson intervals

    Rows
    reference fixtures
    Original Python
    n/a (new in 2026)
    TypeScript test
    stats.test.ts (vs Python and SciPy)
  • Bootstrap, Wilson, kappa, McNemar, OLS; word-list stability

    Rows
    reference fixtures
    Original Python
    n/a (new in 2026)
    TypeScript test
    stats-extra.test.ts (vs numpy, SciPy, statsmodels, scikit-learn, a Python port)
  • Reading-grade intervals, paired gap and debate trends

    Rows
    7,556 documents, 179 debates
    Original Python
    n/a (new in 2026)
    TypeScript test
    readability-stats.test.ts (vs numpy on analytics.db)

“Identical” means every field compared equal as strings: 0 mismatches across 9 document fields and 6 debate fields. The TypeScript suites assert the same with zero tolerance (floating-point values to 1e-9). The last four rows are code written for the revival, so there is no notebook to compare with; their TypeScript is checked against the revival's own Python build and the Python scientific stack instead (scripts/stats_reference.py).

03

Date normalisation

Three functions handle dates, and they disagree on purpose. The listing scrapers turn the ISO content attribute into “September 29, 2024” with strftime("%B %d, %Y"), keeping the wall-clock date of the timestamp rather than converting time zones. The content extractor goes the other way: an ISO string containing both “T” and “+” is kept as ISO, and anything else is tried against four formats in order (%B %d, %Y, %m/%d/%Y, %Y-%m-%d, %d %B %Y) before being passed through unchanged. A trailing “Z” therefore survives in the content phase but not in the listing phase.

Documents, listing phase

"September 29, 2024"

Documents, content phase

"2024-09-29T00:00:00+00:00"

Debates, listing

"September 29, 2024"

04

Transcript splitting

For each paragraph whose first bold element starts with “participants” or “moderators”, the notebook takes the paragraph's HTML, removes a leading <b>LABEL:</b>, replaces each <br> with a sentinel, strips tags, splits on the sentinel and then on semicolons, and trims trailing full stops. That is why names keep a trailing “and”. The plain text is every text node, stripped and joined with newlines, with runs of three or more newlines collapsed. The TypeScript port does the same over an htmlparser2 tree, including BeautifulSoup's serialisation (<br/>, minimal escaping) and Python's definition of whitespace. One known gap: malformed HTML (an unclosed or nested <p>) is repaired differently by htmlparser2 than by Python's html.parser, so a pasted fragment like that can split differently in the playground. All 179 stored transcripts are well formed and match exactly.

A made-up example. Paste the HTML of any transcript from the archive to see what the 2025 notebook would have stored.

Participants

Governor Avery Example (X) and; Senator Blake Sample (Y)

Moderators

Casey Host (Network One); and; Drew Anchor (Network Two)

Participants_List

['Governor Avery Example (X) and', 'Senator Blake Sample (Y)']

Moderators_List

['Casey Host (Network One)', 'and', 'Drew Anchor (Network Two)']

Debate_Content_Text

PARTICIPANTS: Governor Avery Example (X) and Senator Blake Sample (Y) MODERATORS: Casey Host (Network One); and Drew Anchor (Network Two) HOST: Good evening, and welcome to tonight's debate. EXAMPLE : Thank you. It is good to be here & to see you all. SAMPLE: Thank you, Casey. (APPLAUSE)

05

Whose words count

Documents are filed under one candidate, but interviews and town halls also contain interviewers, moderators and audience members. Before counting words, each stored document is split into its paragraphs and read in order. A paragraph that opens with an upper-case label (“ROGAN:”, “AUDIENCE MEMBER:”) switches the current speaker until the next label. Labels that contain the candidate's surname or an office they held (“THE VICE PRESIDENT:”) keep the text; labels that are fields of a press release (“FACT:”, “TIME:”, “NARRATOR:”) are not speaker changes; any other label drops the text that follows. Stage directions such as “(Applause.)” are removed.

This keeps 94.5% of all words and removes 265,530 words spoken by other people. The original word counts are untouched and are what the Explorer reports; the cleaned text feeds the reading grades, the word index and the timelines.

06

Reading grade

The Flesch-Kincaid grade level combines average sentence length and average syllables per word. Words are runs of letters (with internal apostrophes); sentences end at “.”, “!” or “?” followed by a space or quote, after protecting titles (“Mr.”) and initials (“J.D.”); syllables use a standard vowel-group heuristic. Texts under 100 words get no grade. Transcribed speech and press releases are punctuated by different people, so grades compare like with like only roughly, and they say nothing about the quality of what was said.

Flesch-Kincaid grade level
grade = 0.39 × (words ÷ sentences) + 11.8 × (syllables ÷ words) − 15.59

A label keeps its paragraphs if it contains the speaker's surname (here the last word of this box) or an office such as “THE PRESIDENT:”; any other label drops the paragraphs that follow it.

Kept paragraphs

3 kept · 1 dropped

Words

127

Sentences

9

Syllables

184

Grade

7.01

07

Distinctive words: Fightin' Words

The comparison uses the weighted log-odds ratio with an informative Dirichlet prior from Monroe, Colaresi and Quinn (2008), “Fightin' Words: Lexical Feature Selection and Evaluation for Identifying the Content of Political Conflict”, Political Analysis 16(4). The prior is the whole archive's word distribution scaled to α₀ pseudo-words (10,000 by default): it pulls rare words towards the archive-wide rate so that a word used twice by one group and never by the other does not top the list.

The index holds 16,877 words: lower-cased tokens of two or more letters, present in at least five documents, minus a standard English stop-word list. Counts per document are stored as compressed postings; group totals are summed on the server for each request.

For word w, groups A and B
a_w = α₀ × (archive count of w ÷ archive total) δ_w = log[(y_w^A + a_w) / (n^A + α₀ − y_w^A − a_w)] − log[(y_w^B + a_w) / (n^B + α₀ − y_w^B − a_w)] σ²(δ_w) ≈ 1/(y_w^A + a_w) + 1/(y_w^B + a_w) z_w = δ_w / σ(δ_w)

08

Rates and intervals

A period's rate is mentions ÷ words × 10,000, where words are the cleaned own-voice tokens of every document in the period. Treating the count k as Poisson, the exact (Garwood) 95% interval comes from gamma quantiles, computed in TypeScript with a regularised incomplete gamma function and checked against SciPy.

The topic list has 44 entries chosen to cover policy areas that campaigns of both major parties talk about, and both party names. Topic words are matched as whole words, ignoring case, with one exception. The two party topics are defined the same way: the party's noun and its party adjective, counted only when capitalised (“Democrat”, “Democrats”, “Democratic”; “Republican”, “Republicans”). Capitalisation separates “Democratic Party” and “Republican nominee” from the generic words “democratic” and “republican”, which are left out for both. Topics: Economy, Jobs, Inflation, Taxes, Wages, Middle class, Debt and deficit, Trade and tariffs, Manufacturing, Energy, Climate, Health care, Medicare, Social Security, Education, Housing, Infrastructure, Immigration, Border, Crime, Police, Guns, Abortion, Opioids and fentanyl, COVID-19 and pandemic, Veterans, Military, Terrorism, China, Russia, Ukraine, Israel, Iran, Democracy, Freedom and liberty, Constitution, Supreme Court, Corruption, Families, Women, Children, Faith, Democrats (party), Republicans (party).

Exact Poisson interval for k mentions in N words
lower = Γ⁻¹(0.025; k) (0 when k = 0) upper = Γ⁻¹(0.975; k + 1) rate = k / N × 10,000 interval scaled the same way

09

Debate turns and roles

Transcripts from six decades mark speakers in different ways, so a new turn starts at a bold or italic label (“WALZ:”, “The President.”), a plain upper-case label ending in a colon (“MR. NIXON:”), or, in transcripts that use neither, a short name ending in a full stop (“Mr. Newman.”). Lines without a label continue the current speaker. Labels are reduced to a surname; “THE PRESIDENT” and “THE VICE PRESIDENT” become the office holder of that year.

A speaker is a candidate if the surname is on the list below of people who took part as candidates in that cycle's debates (public record), belongs to the party holding the debate when it is a primary, and is named in the page's Participants block when there is one. Every other named speaker is grouped as a moderator, panellist or questioner; audience members, unidentified voices and recorded clips form a third group. Word share stands in for talk time.

Recorded material played during a debate is not live speech. Turns between “[begin video clip]” and “[end video clip]” (and the transcripts' other spellings of these markers), turns tagged “(from videotape.)”, and labels such as “VIDEO CLIP OF …” are counted as Recorded clips, so a candidate heard only in a clip, such as a president quoted at the other party's primary, is not listed as taking part. A clip whose end marker is missing covers only the next speaker. A bold “MODERATOR:” label is a turn unless it heads the list of names at the top of the page, and the build stops if moderators get less than 2% of a debate's words, which would mean their labels were missed. Where a source transcript leaves out a label, the words go to the previous speaker; one such gap in the January 2004 Greenville debate is visible in its moderator share.

Candidates by cycle
1960
Kennedy, Nixon
1976
Carter, Dole, Ford, Mondale
1980
Anderson, Carter, Reagan
1984
Bush, Ferraro, Hart, Jackson, Mondale, Reagan
1988
Bentsen, Bush, Dukakis, Quayle
1992
Bush, Clinton, Gore, Perot, Quayle, Stockdale
1996
Alexander, Buchanan, Clinton, Dole, Dornan, Forbes, Gore, Gramm, Kemp, Keyes, Lugar, Taylor
2000
Bauer, Bradley, Bush, Cheney, Forbes, Gore, Hatch, Keyes, Lieberman, McCain
2004
Bush, Cheney, Clark, Dean, Edwards, Kerry, Kucinich, Lieberman, Sharpton
2008
Biden, Brownback, Clinton, Cox, Dodd, Edwards, Gilmore, Giuliani, Gravel, Huckabee, Hunter, Keyes, Kucinich, McCain, Obama, Palin, Paul, Richardson, Romney, Tancredo, Thompson
2012
Bachmann, Biden, Cain, Gingrich, Huntsman, Johnson, Obama, Paul, Pawlenty, Perry, Romney, Ryan, Santorum
2016
Bush, Carson, Chafee, Christie, Clinton, Cruz, Fiorina, Gilmore, Graham, Huckabee, Jindal, Kaine, Kasich, O'Malley, Pataki, Paul, Pence, Perry, Rubio, Sanders, Santorum, Trump, Walker, Webb
2020
Bennet, Biden, Bloomberg, Booker, Bullock, Buttigieg, Castro, Delaney, Gabbard, Gillibrand, Harris, Hickenlooper, Inslee, Klobuchar, O'Rourke, Pence, Ryan, Sanders, Steyer, Swalwell, Trump, Warren, Williamson, Yang, de Blasio
2024
Biden, Burgum, Christie, DeSantis, Haley, Harris, Hutchinson, Pence, Ramaswamy, Scott, Trump, Vance, Walz

10

Source, licence and limits

Source. Gerhard Peters and John T. Woolley, The American Presidency Project, University of California, Santa Barbara. Copyright © The American Presidency Project. Document and debate pages are linked, never re-published. The site stores titles, dates, counts and short quotations of at most 25 words, each linked to its page.

Coverage. The documents are what the archive files under “campaign documents” between 2016-01-01 and 2024-11-06; that is not everything a campaign published, and the volume per candidate reflects both the campaign and the archive. 26 listing rows were duplicates and are counted once; 46 documents have no text.

Neutrality. Every speaker is processed by the same code. Colours identify the groups being compared, never parties. No measure here is a sentiment, quality or truthfulness score, and nothing predicts an outcome.

Code. MIT licence, covering the code only. The original notebooks and CSVs are kept unchanged in original/ of the GitHub repository.

11

Stability of the word lists

The Fightin' Words z-score treats each word token as an independent draw. Campaign text is clustered: one press release can repeat a county's name thirty times. Every comparison on Distinctive words is therefore checked by resampling documents, not words: the documents of each group are drawn with replacement (200 times, seed 20261010, group A first, then group B, from one mulberry32 stream), every z-score is recomputed with the same prior and α₀, and for each listed word the page reports the share of resamples in which it stays in its side's top 30, the 2.5th to 97.5th percentile of its z-score, the documents that use it and the share of its uses from its heaviest document.

With 200 resamples a kept share has a Monte Carlo standard error of at most 3.5 percentage points. Separately, a comparison tests every word used by either group (up to 16,877), and at |z| = 1.96 about 5% of the words tested would cross the line even if the groups did not differ; the page shows that figure for each comparison. The informative prior makes it a rough guide, but the point stands: the lists are rankings to explore, not a set of findings. Rationale and results: DR-003.

12

Uncertainty in reading grades

The rows behind Readability are not independent. A campaign's documents share writers and a house style, and the debates of one election cycle share candidates and a transcription source. Resampling single documents or debates would treat them as independent and give intervals that are too narrow, so the intervals resample clusters (a cluster bootstrap, 10,000 resamples, seed 20261010): for the mean grade of a cycle and kind of text, whole speakers with all their documents; for a debate trend (the slope of a least-squares line per decade, and a pointwise band for the fitted line), whole cycles with all their debates. Cells with fewer than 5 documents or 10 speakers get a point estimate and no interval: with fewer speakers the cluster bootstrap runs narrow (DR-008).

The difference is not small. For 2024 written releases (1,395 documents from 10 speakers) the speaker-level interval is 10.00 to 12.92, several times wider than resampling documents would suggest. For the primary-debate trend (130 debates in 9 cycles) the cycle-level interval is −0.89 to −0.17 grade levels per decade, against −0.65 to −0.26 if debates were independent. With only 9 to 14 cycles, a cluster bootstrap is itself approximate and can still run a little narrow. A cycle's own mean in the table resamples that cycle's debates and describes that cycle only.

Because the grade is linear in words per sentence (WPS) and syllables per word (SPW), a difference of mean grades splits exactly into a sentence-length part and a word-length part. For the 14 speakers with enough of both kinds of text, each speaker is one paired difference, and the mean gap is −5.5 grade levels (95% t interval −6.9 to −4.1; at n = 14 a percentile bootstrap gives the slightly narrower −6.7 to −4.2), of which −3.2 comes from sentence length. 14 of 14 speakers grade lower when transcribed (exact sign test p = 0.0001). These intervals cover sampling, not measurement: they do not include the effect of who transcribed a debate, which the page shows separately with two events the archive holds in two transcripts. The choice of clustered intervals is recorded in DR-006.

Splitting a difference in mean grade
Δgrade = 0.39 × ΔWPS + 11.8 × ΔSPW (sentence length) (word length)

13

Topic labels: evaluation design

Question. On one-sentence campaign excerpts, how often does a language model give the same policy-topic label as a careful coder, compared with a transparent keyword dictionary?

Set. 120 sentences of 12 to 25 words, 40 per cycle, one per randomly drawn document (seed 20261010), after the quotation rules. The draw is blind to keywords, so it does not favour the dictionary. Gold labels: AI draft, not labelled by a person, so every score is provisional. The gold labels were drafted by the AI coding assistant (Claude) that built the 2026 upgrade, as a single annotator, and no person has labelled them. Treat every score as provisional. Because the draft came from the same model family as the default labeller, agreement with Claude models may be flattered. The same assistant also wrote the keyword dictionary shortly before, so the gold labels are not independent of the baseline either, which could move its scores in either direction. The labels will be replaced by a blind relabel by people who have not seen the draft, not by a review of this draft.

Labellers. The keyword rules, keywords-v1 (10 October 2026), were written before the set was drawn and then frozen. The model sees the codebook (21 CAP-style topics and “none”), five coding rules and up to ten excerpts with opaque ids per request; the system prompt is 4,284 characters. Defaults: Claude Haiku 4.5 at temperature 0, or Claude Sonnet 5.5 at low effort, or an OpenAI model (gpt-5-mini by default).

Metrics. Agreement with gold (Wilson interval), Cohen's kappa (percentile bootstrap over excerpts), agreement on the excerpts with a policy topic, and a paired comparison on the same excerpts: the difference in agreement and in kappa (paired bootstrap, so both labellers see the same resamples) and McNemar's exact test. The keyword baseline on all 120 excerpts: 75.8% agreement (95% CI 67.4 to 82.6), kappa 0.59 (0.45 to 0.71).

14

Assumptions and limitations

  • Coverage. The documents are what the archive files as campaign documents; volumes per candidate reflect the campaign and the archive, not how much anyone said.
  • Own voice. Text is attributed by speaker labels: turns labelled with anyone other than the document's own speaker are dropped, and interviewer text without a label stays in. Debate roles depend on curated candidate lists.
  • Independence. z-scores and Poisson intervals treat word tokens as independent. The word lists now carry a document-level bootstrap check; the timeline intervals do not yet, so they are too narrow when a few documents repeat a term.
  • Readability. Flesch-Kincaid was built for edited English prose. Transcripts are punctuated by transcribers, so grades compare like with like only roughly, and never measure quality.
  • Intervals. Every interval covers sampling variability given the pipeline's choices (tokeniser, stop words, cleaning, roles). None of them covers uncertainty in those choices. Readability intervals resample speakers or election cycles, but with 9 to 24 clusters they are approximate; the topic-label intervals treat the 120 excerpts, one per document, as independent.
  • Topic labels. A small gold set (AI draft, not labelled by a person); each policy topic has zero to 8 excerpts, and “Transportation”, “Social welfare” and “Culture and the arts” have none, so the evaluation says nothing about them; one sentence out of context is hard for any coder. The same AI assistant drafted the gold labels and wrote the keyword dictionary, so the gold set is independent of neither labeller until people relabel it blind.
  • Many tests. Distinctive words tests every indexed word at once; some extreme z-scores are chance.

15

AI use statement

What AI does here. One optional feature: on Topic labels, a language model you choose labels the policy topic of short excerpts so that it can be compared with keyword rules and gold labels. It runs only when you start it, with your own API key.

What it never does. It produces no number anywhere else on the site; every other page is computed without AI. It is never used to compare candidates or parties, and its labels are never presented as facts: every output carries an “AI-generated” label, and the simulated demo is labelled as simulated. A request that fails (a rejected key, a network or rate-limit error, or pressing Stop) is left out of the scores, never counted as the model's wrong answer.

Data sent to the provider. The codebook, the coding rules and the excerpts (one sentence of 25 words or fewer each, with an opaque id), from your browser straight to Anthropic or OpenAI. The speaker, date, title and link are not sent as metadata, but 35 of the 120 excerpts name the candidate in the text, and some name journalists or officials, so the model can often tell whose campaign wrote them. Your key is kept in your browser (sessionStorage, or localStorage if you tick “remember on this device”), never sent to this site's server and never logged. The provider's own terms apply to what you send it.

Human in the loop and audit. You can accept, correct or reject each run, and accept or reject any single call later from the log; corrections are recorded as edits and never change the scores, which always use the model's own labels. Each decision is added to the call's decision history, so a change of mind never erases the earlier decision, and each record holds the generation settings the request was sent with. Every call, failure and simulated run is logged in your browser's IndexedDB with the prompt, the answer, latency, token use and your decision, viewable and exportable (JSON or CSV) on the AI audit log.

Frameworks. This design is informed by the Australian Government's policy for the responsible use of AI in government, the transparency principles of the EU AI Act and the NIST AI Risk Management Framework. It does not claim compliance with any of them.

AI in building the site. The 2026 upgrade was built with an AI coding assistant, which also wrote the keyword dictionary and prepared the draft gold labels (see DR-007, DR-004 and DR-005). Every statistic is computed by code that is tested against independent Python implementations.

16

Model card

The two topic labellers (keyword rules and the bring-your-own-key LLM) have a model card: intended use, data provenance, evaluation with intervals, known failure modes with examples from the gold set, and ethical considerations.

Read the model card →

17

Decision records

Each record states the decision first, then the context, the options, why, what happened (weak numbers included) and what I'd change. Records are never edited; a change of mind gets a new record.
  1. DR-001 · Accepted · 2026 (revival)Show derived statistics and short linked quotations, never the full textsThe website publishes counts, metadata, derived statistics and quotations of 25 words or fewer, each linked to its page on The American Presidency Project. It never serves the full text of a document or a transcript.
  2. DR-002 · Accepted · 2026 (revival)One set of neutral presentation rules for every speaker and every pageEvery speaker goes through the same code, colours mark the groups being compared and never parties, default views compare election cycles rather than people, nothing is scored for sentiment, quality or truth, nothing is predicted, and quotations that name another candidate or use a charged word are skipped.
  3. DR-003 · Accepted · amended by DR-006 · 2026 (revival); bootstrap check 10 October 2026Fightin' Words with an informative prior, exact Poisson rates, and bootstrap checksCompare vocabularies with the weighted log-odds ratio and informative Dirichlet prior of Monroe, Colaresi and Quinn (2008), report term rates per 10,000 words with exact Poisson intervals, and, since the October 2026 upgrade, check every pair of top-word lists by resampling documents.
  4. DR-004 · Accepted · amended by DR-005, DR-007 · 10 October 2026LLM topic labels only as a bring-your-own-key evaluation against keyword rulesOffer language-model topic labels only inside an evaluation: the model labels the same fixed gold set as a frozen keyword dictionary, both are scored with agreement and Cohen's kappa with intervals, calls go straight from the visitor's browser with their own key, only short excerpts are sent, every call is logged in the browser, and a simulated run shows the harness without a key.
  5. DR-005 · Accepted · 10 October 2026Score only the excerpts the model answered forIn the topic-label comparison, score an excerpt only when the model answered for it, leave out excerpts whose request failed or was stopped (for both labellers, so the comparison stays paired), mark such a run incomplete with no verdict, and disclose that the gold labels are independent of neither labeller. This amends the scoring rule of DR-004, which counted every unlabelled excerpt as a wrong answer.
  6. DR-006 · Accepted · amended by DR-008 · 10 October 2026Clustered intervals for readability, and two corrections to DR-003Readability intervals resample clusters rather than rows: whole speakers (with all their documents) for the mean grade of a cycle and kind of text, and whole election cycles (with all their debates) for the debate trends. The 14-speaker paired gap is reported with a t interval and an exact sign test instead of a percentile bootstrap. Two uncertainty statements in DR-003 are corrected here rather than edited there.
  7. DR-007 · Accepted · 10 October 2026Record how the gold set was made, and finish it with a blind relabelRecord how the topic gold labels were made as fields (method, coders, human coders, whether the coders were blind to the keyword dictionary, agreement between coders and agreement with the AI draft), word the provenance note on /topics and /methods from those fields and always show it, and replace the AI-drafted labels with a blind relabel by people, never with a review of the draft in place. This amends the gold-set entry of DR-004, which marked the labels with a single status.
  8. DR-008 · Accepted · 10 October 2026Fewer readability intervals, and speaker-matched twin transcriptsGive a document cell on /readability an interval only when its documents come from at least ten speakers, and compare a speaker across two transcripts of the same event only when both transcripts give that speaker within 2% of the same number of words. This amends DR-006, whose document cells needed only five speakers.

18

What I'd change

  • Have two people label the topic gold set independently, report their agreement and resolve disagreements; grow the set so every topic has at least ten excerpts.
  • Extend the document-level bootstrap (or a negative binomial model) to the timeline, whose Poisson intervals ignore overdispersion.
  • Re-derive reading grades from sentence boundaries normalised across transcript sources, then measure how much of the debate trend survives.
  • Commit dated reference LLM runs, repeated, once there is a small budget, so the comparison is visible without a key and run-to-run variation is measured.
  • Ask the archive's editors how they would like the collected full texts kept, and move them out of the public repository if they prefer.