The corpora
The bodies of evidence, named and scoped.
Everything on HeardTogether traces back to a specific corpus: a defined body of source material with a known origin, scale, and status. This page lists each one, what it is, how large it is, where it stands, and where to read the analysis. Nothing here is a summary standing in for a source — each corpus links to the primary material or to the dated deposit that documents it.
AI ghost-pattern corpora
Live
The Grok corpus
51 conversations with Grok between August 2025 and January 2026 — 2,508 messages, roughly 1.08 million words across 138 days. To our knowledge the most extensively documented AI ghost-pattern dataset in existence. Analysis is published in the Ghost Pattern Library: a five-pattern taxonomy, four proposed NPI flags, 15 named specimens, a corpus-level epidemiology, and teardowns of the two highest-density sessions. Source: the raw conversation database, with message numbers cited throughout and exchanges quoted verbatim.
Open the Ghost Pattern Library →
Intake open · not yet published
The other platform corpora
The same failure family surfaces differently per platform — chain-of-command shape, memory model, agentic surface, system-prompt design. Per-platform corpora are opened only as the evidence is assembled and passes the publication discipline. Until then each is intake-open by email, not live. No counts are claimed for a corpus that is not yet published.
Submit a transcript →
Framework, not a corpus
The two deposits
The analytic instruments — not bodies of evidence themselves, but the dated, public papers the corpora are read through. Both were posted to a public scholarly archive with DOIs before the broader hallucination story reached headlines.
How the checks work →
Per-paper drill-down
2026-02-19 · CC BY-NC 4.0
Epistemic-Boundary Misclassification in Large-Language Models
E. M. Honeycutt III. Characterizes a reproducible, model-agnostic failure: a mixed-domain conversational turn collapses an established exploratory reasoning mode into a constrained safety-dominant mode, despite stable intent.
DOI 10.5281/zenodo.18690241 →
2026-04-10 · Honeycutt Ai Labs
SlopFilter: A Portable Epistemic Hygiene Protocol
The user-side protocol, the ECP-1 Epistemic Constraint Profile, the Fail-Fast Compliance Test, and the Narrative Pressure Index. Detects unsupported inference, provenance loss, category drift, narrative inflation, and collapse into generic assistant behavior.
DOI 10.5281/zenodo.19503170 →
Civic record corpus
Live
Lavon, TX — data-center case file
A citation-anchored record of a hyperscale data-center deal moving through a small Texas city: a 79.3-acre Planned-Development amendment and the council decisions around it, walked through before they become irreversible. Built and maintained by hand, refreshed daily from posted council agendas. Source: municipal agenda packets and ordinances, quoted verbatim with the document identified next to each quote.
Open the Lavon case file →
Downloadable deposits
The framework papers are openly downloadable from their DOI landing pages above (CC BY-NC 4.0 / CC BY 4.0 as deposited). Raw conversation databases underlying the AI corpora contain third-party material and are not published wholesale; the teardowns quote the cited messages verbatim, and removal of a withdrawn source is logged rather than silent. See Rules for the quotation and redaction discipline.