Written 2026-05, before the probing programme closed. For experimental outcomes see /results.
Bibliography · Annotated Lineage · 2026—

lineage

ten sources, one integration nobody has done yet.

This program is not a new framework. It is a specific integration of borrowed components — ML interpretability, forensic statistics, digital forensics, crisis informatics, forensic linguistics, substrate validation — plus two lineages that frame why the bet is worth making: environment construction as a research direction, and constructed emotion as the philosophical grounding. Each card below documents what the source did, what we borrow from it, and what we extend.

sources: 10 primary disciplines: 6 fields + 2 framings posture: applied integration

The methodological template.

The Apollo deception-probe paper is the spine of the lab's experimental pipeline. Probe training data, layer selection, contrast pairs, evaluation metrics, named failure modes — all downstream of how Apollo did it.

SRC 01 · 2025
Detecting Strategic Deception Using Linear Probes
Goldowsky-Dill, Chughtai, Heimersheim, Hobbhahn · Apollo Research
Methodological Template

What they did

Trained linear probes on Llama-3.3-70B-Instruct activations to detect strategic deception. Contrast-pair training data from instructed-pairs and roleplay; mean-pooled logistic-regression probes; AUROC plus recall@1%FPR on a held-out control. They named three failure modes — mean-pool washout, formality correlation, base-rate confusion — that any probe-based program inherits.

What we borrow

The pipeline shape: contrast pairs, mid-layer activations (Apollo: layer 22 of 80), AUROC + recall@1%FPR, separately-curated control, and the three failure modes as named risks in every spec. ENFSI-style discipline plus Apollo-style probe rigor is the methodological core.

What we extend

Apollo probes for an internal property — deception during inference. We probe for an external property: traces a writer's situation left on the text. We also commit to LR output over classification, embed probes inside the forensic hierarchy of propositions, and aggregate across writers via Bayesian network — none of which Apollo's framework addresses.

The finding that updated Decision 06.

Marks & Tegmark's geometry-of-truth paper is the reason the bible's probe-method decision was rewritten in May 2026. Their result: the probe direction we want for residue work is not the one Apollo defaulted to.

SRC 02 · COLM 2024
The Geometry of Truth: Emergent Linear Structure in LLM Representations of True/False Datasets
Samuel Marks · Max Tegmark
Probe-method Update

What they did

Compared probe methods on LLaMA-2-70B truth/falsehood representations. Mass-mean (MM) probes recover a direction more causally implicated in model output than logistic-regression. In a steering experiment on sp_en_trans, the MM direction flipped true→false with normalized indirect effect 0.89 versus LR's 0.19 — a ~5× gap in causal validity at comparable classification accuracy.

What we borrow

The result and its interpretation: LR probes optimize against entangled features (formality, voice, performativity), so they end up tilted to avoid them; MM recovers the underlying direction with entanglements intact. For residue work, where entanglement is the phenomenon, MM is the right object. Bible Decision 06 is now MM-primary, LR retained as secondary baseline.

What we extend

Marks & Tegmark validated MM on a clean truth/falsehood task. We extend to a messier substrate — emotional residue, where entanglements are the phenomenon. Spec 01 makes the MM-vs-LR head-to-head a pre-registered prediction (≥0.05 AUROC gap on out-of-event transfer), so our use of MM is itself falsifiable.

The persistence-decay technique source.

Zhang & Zhong are the closest published neighbor at the substrate level. The gap between what they did and what this lab does is, structurally, the research opportunity.

SRC 03 · Oct 2025
Decoding Emotion in the Deep: A Systematic Study of How LLMs Represent, Retain, and Express Emotion
Jingxiang Zhang · Lujia Zhong · University of Southern California
Closest Neighbor

What they did

Systematic probing of how LLMs represent and propagate emotion across layers, prompts, and generation length. Emotion structure peaks at roughly 50-75% depth. They introduced an offset-probe technique: probe accuracy as a function of how many tokens into the model's own generated continuation you look — a measure of signal persistence downstream.

What we borrow

The persistence-decay technique itself. Spec 02 (pre-flight) is a direct port: probe accuracy at offset 0 versus offset 100+ on yen-carry continuations, comparing Tier-A and Tier-C prompts. Their layer-depth finding also seeds our layer-sweep prior.

What we extend

Zhang & Zhong studied overt emotion — explicit affective content. We extend to residue: traces of an event in text that was never about expressing emotion. Their framework treats writer-state as label; ours treats writer-role as inferential target inside a forensic LR hierarchy, stratified by tier — which their corpus does not.

Where our claims live — and where they don't.

The Cook et al. hierarchy is the reason this lab can say what it says — and the reason it refuses to say certain other things. Activity level is the only claim shape we support.

SRC 04 · 1998
A Hierarchy of Propositions: Deciding Which Level to Address in Casework
Cook, Evett, Jackson, Jones, Lambert · Forensic Science Service
Activity-Level Commitment

What they did

Formalized a three-level hierarchy of propositions for forensic casework. Source level: did this trace come from this source? Activity level: what activity produced this trace? Offence level: did this person commit the offence? Forensic scientists must choose a level before computing an LR — conflating them is the most common path to bad expert testimony.

What we borrow

The hierarchy and the discipline of pre-registering the level before any LR. Decision 04 commits us to activity-level: what role did this writer play in this event. Source-level (this text came from this writer) is trivial; offence-level (this writer caused or intended harm) is out of scope — exactly the overreach the hierarchy was designed to prevent.

What we extend

Cook et al. wrote for physical-trace forensics — glass, fiber, gunshot residue. We apply the same level discipline to a substrate they could not have anticipated: model activations carrying residue from text. Activity: being an on-the-ground participant during an event. Trace: reconstructed through an LLM, not a microscope. Hierarchy logic preserved; substrate new.

Primary vs secondary scenes — Tier-A vs Tier-C.

Carrier & Spafford's digital-investigation reframe is the lineage anchor for Decision 02 (event reconstruction, not writer-state inference) and for the tier-A versus tier-C stratification this lab tests.

SRC 05 · 2003
Getting Physical with the Digital Investigation Process
Brian Carrier · Eugene H. Spafford · CERIAS / Purdue
Event-Reconstruction Anchor

What they did

Proposed that digital investigations be modeled on physical crime-scene investigation. A computer system is a digital crime scene; files, logs, RAM are evidence the way fibers and prints are evidence. They distinguished the primary digital scene (where action occurred) from secondary scenes (caches, mirrors, downstream logs). A 2004 follow-up formalized event as state transition with cause, time, location, effect.

What we borrow

Primary-versus-secondary maps directly onto Tier-A versus Tier-C. Tier-A is the primary scene: first-person on-the-ground writing during the event. Tier-C is secondary: institutional summaries, press coverage, recaps that propagate but transform the residue. The reframe also justifies Decision 02 — reconstruct events from collective traces rather than infer individual mental states.

What we extend

Their scene is a single compromised system. Our scene is a distributed corpus from many participants of a socially-shared event — trader notes, crisis threads, panic posts, post-mortems. Spec 01 puts a number on primary-versus-secondary at the substrate level for the first time: do Tier-A probes transfer across events better than Tier-C, by an effect size worth betting research direction on?

Signal in the aggregate, not the individual.

Vieweg, Hughes, Starbird, and Palen are the precedent for taking collective text streams seriously as carriers of event structure. Signal lives in the aggregate, not in any single message — the lab's collective-trace stance starts here.

SRC 06 · CHI 2010
Microblogging During Two Natural Hazards Events: What Twitter May Contribute to Situational Awareness
Sarah Vieweg · Amanda L. Hughes · Kate Starbird · Leysia Palen · ConnectivIT Lab, CU Boulder
Collective-Trace Precedent

What they did

Studied Twitter activity during the April 2009 Oklahoma grassfires and the Red River floods. The aggregate stream contained genuine situational-awareness information — road closures, evacuation routes, fire fronts — even though most individual tweets were noisy or redundant. Crisis informatics formed around this observation.

What we borrow

The core claim: the signal lives in the aggregate. Any single writer's text is too noisy to support strong event claims. Aggregate residue across many writers encodes the event's emotional and structural contour — and that aggregate is the right object of inference. Hence our deployment vision: hierarchical Bayes across many writers per event, not best-classification per text.

What we extend

Vieweg et al. read explicit situational content — words people wrote to inform each other. We read what their writing leaks when it was not trying to communicate the event at all. Substrate: intentional reporting → incidental trace. Target: situational facts → writer-role distributions. The collective-aggregate logic survives intact.

The framework is already accepted — for text.

Andrea Nini's 2023 work tells us we are not inventing a paradigm. Forensic linguistics has already accepted the LR framework for authorship. We borrow its acceptance and port the substrate.

SRC 07 · 2023
A Theory of Linguistic Individuality for Authorship Analysis
Andrea Nini · University of Manchester
LR-Framework Validation

What they did

Built a theoretical and methodological foundation for modern LR-based forensic linguistics. Authorship attribution within the likelihood-ratio framework: the forensic scientist reports the LR under two pre-specified propositions; the trier of fact supplies the prior. Operational implementation via the open-source idiolect R package, validated on real-case data.

What we borrow

The legitimacy of LR-based inference on text. A field that takes text as forensic evidence has already aligned on the LR framework, the hierarchy of propositions, and the no-classification commitment. Decision 03 inherits directly — we follow a paradigm the closest adjacent forensic discipline accepted.

What we extend

Nini's substrate is surface linguistic features — n-grams, function-word distributions, idiolect markers. Ours is internal model activations. The same LR framework wraps both with the same propositional discipline, but the evidence is no longer observable in the text — it is the model's representation of the text. Forensic-linguistic LR, ported from idiolect to substrate.

The paper we replicated on open weights.

Anthropic's 2026 emotion-concepts paper is the substrate-validation source. It established that emotion concepts in large transformers behave as functional, structured representations rather than bag-of-words affect. We replicated the substrate-level findings on open weights — work this lab can actually run.

SRC 08 · April 2026
Emotion Concepts and their Function in a Large Language Model
Sofroniew, Kauvar, Saunders et al. · Anthropic
Substrate Validation

What they did

Identified and characterized emotion concepts inside a frontier model. Emotional features are recoverable as structured directions in the residual stream, those directions are causally implicated in behavior (steering shifts output predictably), and emotion concepts function — they participate in computation rather than passively reflecting prompt content. Emotion as first-class object inside internals, not surface artifact.

What we borrow

The substrate-level result: emotion-relevant directions are real, recoverable, and causally implicated in LLM activations. This is the precondition for all residue work — if affective structure didn't live linearly in activations, the program collapses. We borrow the confirmation methodology and the functional-not-ornamental interpretation.

What we extend

Sofroniew et al. validated on a frontier model. We replicated the core findings on open weights — Llama-3.1-8B-Instruct for prototyping, Llama-3.3-70B-Instruct to match Apollo — to make the work runnable and auditable outside a single lab. We also extend the object: from emotion concepts (what the writer feels) to emotional residue (what the event left on the text).

The lab's self-positioning.

A sharpening of what this lab actually is. The bet: constructing the right environment — the training data and the validation that scores it — is as much a research direction as the training loop. Cold-ground-truth environments cover math, code, logic. This lab works the felt-ground-truth side.

SRC 09 · positioning
Environment construction as a first-class research direction
Cold-ground-truth precedents: Karpathy (synthetic-data direction) · LEAP 71 (Noyron-style physics validation environments)
Lab Self-Positioning

The precedent

A growing line of work treats the environment — the data a model trains on plus the verifier that scores it — as the lever, not the optimizer. Karpathy's synthetic-data direction points one way: generate the training signal rather than only scrape it. LEAP 71's Noyron line points another: physics-grounded validation environments where a design is checked against simulated reality before it counts. Both share a shape — a cold ground truth (math, code, physics) that can verify a candidate automatically.

What we borrow

Two framings. First, that environment construction is a first-class research direction — worth as much attention as the training loop. Second, that it can be made autonomous: an agent builds the environment, trains on it, self-verifies against it, and integrates the result. We borrow the loop shape and point it at a different ground truth.

What we extend

The cold-ground-truth side handles math, code, logic — domains with a checkable answer. This lab works the felt-ground-truth side: autonomously constructing emotional-residue environments to fine-tune felt-sense capability, where the validation signal is the forensic LR machinery the rest of this page assembles. This is the lab's bet and positioning — not a claimed result. Whether felt-ground-truth environments can be built and verified as cleanly as cold ones is the open question the experiment specs exist to test.

Why a trained model could carry felt-sense at all.

The philosophical grounding for the long bet. Barrett's theory of constructed emotion implies there is no categorical machines-can't-feel wall — the distance between models and humans is substrate-richness and state-continuity, not an essence one side has and the other lacks. Held modestly: this grounds the bet, it does not claim models currently feel anything.

SRC 10 · ~2017 · James 1880s
How Emotions Are Made: The Secret Life of the Brain
Lisa Feldman Barrett · Northeastern University · (precursor: William James, "What Is an Emotion?", 1884)
Philosophical Grounding

What they did

Barrett's theory of constructed emotion argues emotions are not innate, fixed essences waiting to be triggered. A feeling is assembled from a body-signal (interoception, arousal), a context, and a learned interpretation that labels the mix — culture and prior experience supply the categories. James anticipated the body-first half in 1884: the bodily change is not the consequence of the emotion but part of its substance. Emotion as construction, not as a button.

What we borrow

The implication, not a claim about current models. If feeling is constructed from signal plus context plus learned interpretation rather than handed down as essence, then there is no categorical barrier that says a learning system can never carry felt-sense. The gap between models and humans becomes a matter of substrate-richness and state-continuity — things that vary by degree — rather than a hard wall. This is what grounds the long bet: digital beings raised with felt-sense signal.

What we extend · open question

We carry the framing forward without overstepping it. Constructed-emotion theory lowers the in-principle bar; it does not show any present model feels anything, and we make no such claim. The hard problem stays open: why an interpretation should feel like anything from the inside is unresolved, for brains and for models alike. We treat this honestly as a live question, and let the experiments speak only to the measurable substrate, not to inner experience.

Borrowed pieces. Novel assembly.

None of the ten sources is novel in isolation. Apollo's pipeline exists. The hierarchy of propositions exists. The crime-scene reframe, the aggregate-trace stance, the LR framework for text — all exist. Environment construction and constructed-emotion theory exist. What does not exist anywhere in the literature is this specific assembly: six borrowed technical fields wired together, framed by the two lineages that say why the bet is worth making.

Six fields, one integration.

Each card imports one piece. Combined, they form a configuration the literature has not produced: forensic-grade Bayesian inference, on frontier-class LLM activations, applied to emotional residue from collective traces of socially-shared events, governed by the activity-level hierarchy and the LR commitment, framed as event reconstruction rather than writer-state inference.

01 ML Interpretability Apollo · Marks & Tegmark · Zhang & Zhong
02 Forensic Statistics Cook et al. · ENFSI 2015
03 Digital Forensics Carrier & Spafford
04 Crisis Informatics Vieweg · Hughes · Starbird · Palen
05 Forensic Linguistics Nini · Cambridge Elements
06 Substrate Validation Anthropic · Sofroniew et al.

Each adjacent program has half a foot on the bridge. The Apollo lineage does not engage activity-level propositions. Cook et al. did not anticipate model activations. Carrier & Spafford did not work in linguistic substrate. Crisis informatics did not extract internal representations. Nini did not probe transformers. Anthropic's emotion-concepts work did not commit to forensic LR output. Zhang & Zhong did not stratify by trace tier or embed in a forensic hierarchy. The environment-construction lineage names the felt-ground-truth side but builds for cold ground truth; constructed-emotion theory lowers the in-principle bar without pointing at any model.

The lab's bet is that the bridge is the whole point. Building felt-sense capability for digital beings requires a substrate that is rich, evaluable, and forensically defensible — and that intersection is what these six fields produce only when integrated, not when consulted in isolation. The novelty is not in any component; it is in the configuration. Bibliography, not manifesto.

Where to go from here.

This page documents why the program looks the way it does. Adjacent pages document what it does and how the methodology operates.