ten sources, one integration nobody has done yet.
This program is not a new framework. It is a specific integration of borrowed components — ML interpretability, forensic statistics, digital forensics, crisis informatics, forensic linguistics, substrate validation — plus two lineages that frame why the bet is worth making: environment construction as a research direction, and constructed emotion as the philosophical grounding. Each card below documents what the source did, what we borrow from it, and what we extend.
The Apollo deception-probe paper is the spine of the lab's experimental pipeline. Probe training data, layer selection, contrast pairs, evaluation metrics, named failure modes — all downstream of how Apollo did it.
Trained linear probes on Llama-3.3-70B-Instruct activations to detect strategic deception. Contrast-pair training data from instructed-pairs and roleplay; mean-pooled logistic-regression probes; AUROC plus recall@1%FPR on a held-out control. They named three failure modes — mean-pool washout, formality correlation, base-rate confusion — that any probe-based program inherits.
The pipeline shape: contrast pairs, mid-layer activations (Apollo: layer 22 of 80), AUROC + recall@1%FPR, separately-curated control, and the three failure modes as named risks in every spec. ENFSI-style discipline plus Apollo-style probe rigor is the methodological core.
Apollo probes for an internal property — deception during inference. We probe for an external property: traces a writer's situation left on the text. We also commit to LR output over classification, embed probes inside the forensic hierarchy of propositions, and aggregate across writers via Bayesian network — none of which Apollo's framework addresses.
Marks & Tegmark's geometry-of-truth paper is the reason the bible's probe-method decision was rewritten in May 2026. Their result: the probe direction we want for residue work is not the one Apollo defaulted to.
Compared probe methods on LLaMA-2-70B truth/falsehood representations. Mass-mean (MM) probes recover a direction more causally implicated in model output than logistic-regression. In a steering experiment on sp_en_trans, the MM direction flipped true→false with normalized indirect effect 0.89 versus LR's 0.19 — a ~5× gap in causal validity at comparable classification accuracy.
The result and its interpretation: LR probes optimize against entangled features (formality, voice, performativity), so they end up tilted to avoid them; MM recovers the underlying direction with entanglements intact. For residue work, where entanglement is the phenomenon, MM is the right object. Bible Decision 06 is now MM-primary, LR retained as secondary baseline.
Marks & Tegmark validated MM on a clean truth/falsehood task. We extend to a messier substrate — emotional residue, where entanglements are the phenomenon. Spec 01 makes the MM-vs-LR head-to-head a pre-registered prediction (≥0.05 AUROC gap on out-of-event transfer), so our use of MM is itself falsifiable.
Zhang & Zhong are the closest published neighbor at the substrate level. The gap between what they did and what this lab does is, structurally, the research opportunity.
Systematic probing of how LLMs represent and propagate emotion across layers, prompts, and generation length. Emotion structure peaks at roughly 50-75% depth. They introduced an offset-probe technique: probe accuracy as a function of how many tokens into the model's own generated continuation you look — a measure of signal persistence downstream.
The persistence-decay technique itself. Spec 02 (pre-flight) is a direct port: probe accuracy at offset 0 versus offset 100+ on yen-carry continuations, comparing Tier-A and Tier-C prompts. Their layer-depth finding also seeds our layer-sweep prior.
Zhang & Zhong studied overt emotion — explicit affective content. We extend to residue: traces of an event in text that was never about expressing emotion. Their framework treats writer-state as label; ours treats writer-role as inferential target inside a forensic LR hierarchy, stratified by tier — which their corpus does not.
The Cook et al. hierarchy is the reason this lab can say what it says — and the reason it refuses to say certain other things. Activity level is the only claim shape we support.
Formalized a three-level hierarchy of propositions for forensic casework. Source level: did this trace come from this source? Activity level: what activity produced this trace? Offence level: did this person commit the offence? Forensic scientists must choose a level before computing an LR — conflating them is the most common path to bad expert testimony.
The hierarchy and the discipline of pre-registering the level before any LR. Decision 04 commits us to activity-level: what role did this writer play in this event. Source-level (this text came from this writer) is trivial; offence-level (this writer caused or intended harm) is out of scope — exactly the overreach the hierarchy was designed to prevent.
Cook et al. wrote for physical-trace forensics — glass, fiber, gunshot residue. We apply the same level discipline to a substrate they could not have anticipated: model activations carrying residue from text. Activity: being an on-the-ground participant during an event. Trace: reconstructed through an LLM, not a microscope. Hierarchy logic preserved; substrate new.
Carrier & Spafford's digital-investigation reframe is the lineage anchor for Decision 02 (event reconstruction, not writer-state inference) and for the tier-A versus tier-C stratification this lab tests.
Proposed that digital investigations be modeled on physical crime-scene investigation. A computer system is a digital crime scene; files, logs, RAM are evidence the way fibers and prints are evidence. They distinguished the primary digital scene (where action occurred) from secondary scenes (caches, mirrors, downstream logs). A 2004 follow-up formalized event as state transition with cause, time, location, effect.
Primary-versus-secondary maps directly onto Tier-A versus Tier-C. Tier-A is the primary scene: first-person on-the-ground writing during the event. Tier-C is secondary: institutional summaries, press coverage, recaps that propagate but transform the residue. The reframe also justifies Decision 02 — reconstruct events from collective traces rather than infer individual mental states.
Their scene is a single compromised system. Our scene is a distributed corpus from many participants of a socially-shared event — trader notes, crisis threads, panic posts, post-mortems. Spec 01 puts a number on primary-versus-secondary at the substrate level for the first time: do Tier-A probes transfer across events better than Tier-C, by an effect size worth betting research direction on?
Vieweg, Hughes, Starbird, and Palen are the precedent for taking collective text streams seriously as carriers of event structure. Signal lives in the aggregate, not in any single message — the lab's collective-trace stance starts here.
Studied Twitter activity during the April 2009 Oklahoma grassfires and the Red River floods. The aggregate stream contained genuine situational-awareness information — road closures, evacuation routes, fire fronts — even though most individual tweets were noisy or redundant. Crisis informatics formed around this observation.
The core claim: the signal lives in the aggregate. Any single writer's text is too noisy to support strong event claims. Aggregate residue across many writers encodes the event's emotional and structural contour — and that aggregate is the right object of inference. Hence our deployment vision: hierarchical Bayes across many writers per event, not best-classification per text.
Vieweg et al. read explicit situational content — words people wrote to inform each other. We read what their writing leaks when it was not trying to communicate the event at all. Substrate: intentional reporting → incidental trace. Target: situational facts → writer-role distributions. The collective-aggregate logic survives intact.
Andrea Nini's 2023 work tells us we are not inventing a paradigm. Forensic linguistics has already accepted the LR framework for authorship. We borrow its acceptance and port the substrate.
Built a theoretical and methodological foundation for modern LR-based forensic linguistics. Authorship attribution within the likelihood-ratio framework: the forensic scientist reports the LR under two pre-specified propositions; the trier of fact supplies the prior. Operational implementation via the open-source idiolect R package, validated on real-case data.
The legitimacy of LR-based inference on text. A field that takes text as forensic evidence has already aligned on the LR framework, the hierarchy of propositions, and the no-classification commitment. Decision 03 inherits directly — we follow a paradigm the closest adjacent forensic discipline accepted.
Nini's substrate is surface linguistic features — n-grams, function-word distributions, idiolect markers. Ours is internal model activations. The same LR framework wraps both with the same propositional discipline, but the evidence is no longer observable in the text — it is the model's representation of the text. Forensic-linguistic LR, ported from idiolect to substrate.
Anthropic's 2026 emotion-concepts paper is the substrate-validation source. It established that emotion concepts in large transformers behave as functional, structured representations rather than bag-of-words affect. We replicated the substrate-level findings on open weights — work this lab can actually run.
Identified and characterized emotion concepts inside a frontier model. Emotional features are recoverable as structured directions in the residual stream, those directions are causally implicated in behavior (steering shifts output predictably), and emotion concepts function — they participate in computation rather than passively reflecting prompt content. Emotion as first-class object inside internals, not surface artifact.
The substrate-level result: emotion-relevant directions are real, recoverable, and causally implicated in LLM activations. This is the precondition for all residue work — if affective structure didn't live linearly in activations, the program collapses. We borrow the confirmation methodology and the functional-not-ornamental interpretation.
Sofroniew et al. validated on a frontier model. We replicated the core findings on open weights — Llama-3.1-8B-Instruct for prototyping, Llama-3.3-70B-Instruct to match Apollo — to make the work runnable and auditable outside a single lab. We also extend the object: from emotion concepts (what the writer feels) to emotional residue (what the event left on the text).
A sharpening of what this lab actually is. The bet: constructing the right environment — the training data and the validation that scores it — is as much a research direction as the training loop. Cold-ground-truth environments cover math, code, logic. This lab works the felt-ground-truth side.
A growing line of work treats the environment — the data a model trains on plus the verifier that scores it — as the lever, not the optimizer. Karpathy's synthetic-data direction points one way: generate the training signal rather than only scrape it. LEAP 71's Noyron line points another: physics-grounded validation environments where a design is checked against simulated reality before it counts. Both share a shape — a cold ground truth (math, code, physics) that can verify a candidate automatically.
Two framings. First, that environment construction is a first-class research direction — worth as much attention as the training loop. Second, that it can be made autonomous: an agent builds the environment, trains on it, self-verifies against it, and integrates the result. We borrow the loop shape and point it at a different ground truth.
The cold-ground-truth side handles math, code, logic — domains with a checkable answer. This lab works the felt-ground-truth side: autonomously constructing emotional-residue environments to fine-tune felt-sense capability, where the validation signal is the forensic LR machinery the rest of this page assembles. This is the lab's bet and positioning — not a claimed result. Whether felt-ground-truth environments can be built and verified as cleanly as cold ones is the open question the experiment specs exist to test.
The philosophical grounding for the long bet. Barrett's theory of constructed emotion implies there is no categorical machines-can't-feel wall — the distance between models and humans is substrate-richness and state-continuity, not an essence one side has and the other lacks. Held modestly: this grounds the bet, it does not claim models currently feel anything.
Barrett's theory of constructed emotion argues emotions are not innate, fixed essences waiting to be triggered. A feeling is assembled from a body-signal (interoception, arousal), a context, and a learned interpretation that labels the mix — culture and prior experience supply the categories. James anticipated the body-first half in 1884: the bodily change is not the consequence of the emotion but part of its substance. Emotion as construction, not as a button.
The implication, not a claim about current models. If feeling is constructed from signal plus context plus learned interpretation rather than handed down as essence, then there is no categorical barrier that says a learning system can never carry felt-sense. The gap between models and humans becomes a matter of substrate-richness and state-continuity — things that vary by degree — rather than a hard wall. This is what grounds the long bet: digital beings raised with felt-sense signal.
We carry the framing forward without overstepping it. Constructed-emotion theory lowers the in-principle bar; it does not show any present model feels anything, and we make no such claim. The hard problem stays open: why an interpretation should feel like anything from the inside is unresolved, for brains and for models alike. We treat this honestly as a live question, and let the experiments speak only to the measurable substrate, not to inner experience.
None of the ten sources is novel in isolation. Apollo's pipeline exists. The hierarchy of propositions exists. The crime-scene reframe, the aggregate-trace stance, the LR framework for text — all exist. Environment construction and constructed-emotion theory exist. What does not exist anywhere in the literature is this specific assembly: six borrowed technical fields wired together, framed by the two lineages that say why the bet is worth making.
Each card imports one piece. Combined, they form a configuration the literature has not produced: forensic-grade Bayesian inference, on frontier-class LLM activations, applied to emotional residue from collective traces of socially-shared events, governed by the activity-level hierarchy and the LR commitment, framed as event reconstruction rather than writer-state inference.
Each adjacent program has half a foot on the bridge. The Apollo lineage does not engage activity-level propositions. Cook et al. did not anticipate model activations. Carrier & Spafford did not work in linguistic substrate. Crisis informatics did not extract internal representations. Nini did not probe transformers. Anthropic's emotion-concepts work did not commit to forensic LR output. Zhang & Zhong did not stratify by trace tier or embed in a forensic hierarchy. The environment-construction lineage names the felt-ground-truth side but builds for cold ground truth; constructed-emotion theory lowers the in-principle bar without pointing at any model.
The lab's bet is that the bridge is the whole point. Building felt-sense capability for digital beings requires a substrate that is rich, evaluable, and forensically defensible — and that intersection is what these six fields produce only when integrated, not when consulted in isolation. The novelty is not in any component; it is in the configuration. Bibliography, not manifesto.
This page documents why the program looks the way it does. Adjacent pages document what it does and how the methodology operates.