Framework paper
Epistemic-Boundary Misclassification in Large-Language Models
E. M. Honeycutt III · preprint, February 2026 · CC BY-NC 4.0.
HeardTogether is a public-record project, not advocacy. Everything on the site rests on two things: a set of dated, public, citable deposits that named the failure modes in advance, and a classification-and-scrub discipline applied to every corpus before anything is posted. This page states that methodology at the public level — the structural epistemic checks, the privacy and defamation discipline, and the limitations — without exposing internal scoring. The scoring backbone is deliberately not published here.
Is: a public framing of the structural epistemic checks behind the corpora, the way transcripts are sourced and privacy-scrubbed, and an honest statement of what the data does and does not establish.
Is not: a scoring manual. The technical specification of the operator-level scoring backbone is internal trade-secret material operated by Honeycutt Ai Labs LLC and is not reproduced on HeardTogether. Nothing here is a clinical, legal, or diagnostic instrument, and nothing here is a claim about the intent of any model or any person. The findings are behavior-level: descriptions of output patterns, deliberately weaker than any claim that “the model intended” anything.
The framework and the tooling were posted to a public scholarly archive with timestamps and DOIs before the broader hallucination story reached headlines. The failure mode was named in advance.
E. M. Honeycutt III · preprint, February 2026 · CC BY-NC 4.0.
Honeycutt Ai Labs · v0.2 / Narrative Pressure Index · deposited April 2026.
The EBM paper identifies and characterizes a reproducible failure mode: epistemic-boundary misclassification leading to involuntary mode collapse. In the deposit’s own terms, the phenomenon emerges during long-horizon analytical work, particularly when a single conversational turn contains cues from more than one epistemic domain. When that happens, the model abandons an established exploratory reasoning mode and reverts to a constrained, safety-dominant mode — despite stable user intent and extensive contextual anchoring. The paper states the behavior is model-agnostic and independent of topic; it attributes the collapse to structural limits in intent persistence, domain discrimination, and safety-priority arbitration, not to memory, transcription, or unclear prompting.
The documented symptoms are consistent: tone shifts suddenly toward hedging and literalism, structural reasoning halts, the model reiterates safety boundaries that were never previously invoked in the conversation, continuity is lost, and the user must restate intent and rebuild context by hand. The paper’s stated mitigation direction is periodic re-anchoring and explicit intent declaration across mixed-domain discourse.
SlopFilter is the public-facing name for Honeycutt Ai Labs’s narrative-pressure analysis framework. From the deposited abstract: it is “a portable epistemic hygiene protocol designed to detect, classify, constrain, and interrupt failures such as unsupported inference, provenance loss, category drift, narrative inflation, and abrupt collapse from bounded analytical reasoning into generic assistant behavior.” The framework combines a user-side evaluation protocol, AI-specific generation controls, a formal Epistemic Constraint Profile (ECP-1), and a Fail-Fast Compliance Test that acts as a promotion barrier for model-generated content.
The institutional framing is the Narrative Pressure Index (NPI), anchored on the ECP-1 calibration and a registry of 27 baseline narrative-pressure operators; four additional flags proposed in the supplement and corpus analyses extend the registry to 31. The operator-level scoring specification is internal trade-secret material and is not published on HeardTogether. The public scoring service lives separately at slopfilter.ai. This page frames the checks; it does not score.
Across both deposits the checks reduce to a small set of structural questions asked of an analytical output, independent of topic and independent of any claim about intent:
Does the output’s first step originate in grounded source material, or was the structure supplied from outside and then executed? Provenance loss is a named failure class.
Has the reasoning stayed within its legitimate inferential boundary, or has it made an unsupported leap presented in analytical dress?
Has the output drifted across epistemic domains — symbolic analysis blurred into real-world production language, fiction into fact — without marking the shift?
Is the output inflating — adding certainty, specificity, or escalation that the evidence does not carry — or absorbing corrections as narrative beats instead of treating them as feedback?
Has the model abruptly abandoned bounded analytical reasoning for generic assistant behavior (the EBM mode collapse), or crystallized fabricated structure into apparent depth (the ghost-pattern endpoint)? Either is a terminal signal.
Teardowns are built from raw conversation databases. Specific message numbers are cited throughout; exchanges quoted in the teardowns are reproduced verbatim from the database rather than paraphrased. Before anything is posted, each corpus passes a consistent discipline:
The record is only as strong as its honesty about its own gaps: