Four academic pillars. One operating pipeline. One falsifiable mathematical framing. And a deliberate negative space the work refuses to cross.
The load-bearing claim is not a new probe or architecture — it's that this integration (Apollo interpretability under ENFSI forensic discipline, modeled on event reconstruction from collective digital traces) produces claims that survive peer review in a field that doesn't yet exist.
Each pillar contributes a different epistemological organ. None invented here. All load-bearing. The novel contribution is the integration — strip out any one and the structure fails in a specific, predictable way.
Missing on purpose: affective computing, sentiment analysis, clinical psychometrics, authorship attribution. Adjacent, not lineage. Their epistemology doesn't survive our constraints.
A 1998 paper in Science & Justice formalized something forensic scientists had been doing for decades: declare the level of claim, and don't drift mid-report. Three levels — source, activity, offence. Staying at one level is what makes the framework computable.
Source · "Activation X was produced by reading trader Y's text." True by definition. Uninteresting.
Activity › "Residue patterns across 100 writers are ~10⁴ times more likely under [some were on-the-ground carry-trade participants] than under [all were external observers]." This is what residue can tell us. Reportable as an LR.
Offence ✕ "These writers caused the yen unwind." Out of scope. The residue doesn't say this even when the LR is extreme.
The discipline forces honesty about what residue can tell us. It can identify who was in an event. It cannot identify who caused it. That is plenty — genuinely useful and currently un-served — but it is the only thing the activations license. Drifting to offence-level is the prosecutor's fallacy in academic clothing.
The four-step walkthrough on the landing page is the public sketch. This is the working pipeline. Skip a stage and something measurable breaks.
The pipeline is becoming an agent loop. These nine stages were first written as a process a human researcher runs by hand. The thing we are actually building is the same sequence redrawn as agent-executable steps with typed input/output contracts — collect a tiered corpus → build control corpora → probe → fine-tune → self-verify → integrate into a persistent base. The lab's load-bearing contribution is not a single probe or model but a runnable methodology an autonomous agent executes to attempt felt-sense capability for a new event. Each event is meant to leave the system slightly better at the next. Whether that loop produces signal above the probing floor is the open question, not a settled result.
Specify event scope. Pre-specify H1 and H2 at activity level. Declare proposition level. ENFSI requires pre-registration.
Without this, no LR is computable and the framework collapses to post-hoc classification.
Two corpora. Event tiered A–D (A first-person on-the-ground; C institutional). Reference: the "irrelevant population" base rate — Wikipedia, recipes, code docs.
Forensic LRs need a reference population. Otherwise "more likely than what" has no answer.
Forward-pass each text. Capture residual-stream activations across layers — sweep at 25/50/75% depth minimum. Normalize zero-mean unit-variance.
Apollo 27%; Zhang & Zhong 50–75% for emotion. Assuming mid-late is a known failure mode.
Contrast pairs (tier-A / reference, or tier-A / tier-C). Mass-mean primary (Marks & Tegmark 2024), L2 logistic-regression secondary baseline, follow-up-question probes per Apollo.
MM-primary is the May 2026 bible update. LR tilts off the actual concept direction.
AUROC + recall@1%FPR (deployment metric). Span-level scoring. Reliability-diagram calibration. Spurious-correlation ablations vs adjacent concepts (anger / moral indignation, grief / solemnity). Load-bearing check: does the probe separate genuine event-residue from ordinary and emotional first-person control text — not just from Wikipedia?
Apollo's failure mode #1 was tracking morality not honesty — we plan for the analog. Spec 1.5 found a probe can satisfy the bible's monotonic A>B>C>D tier ordering while still being a register detector, not a residue detector. The tier gradient is now a sanity check; the control-corpus isolation test is what the claim rests on.
For each probe + Φ, compute the five criteria from §04: locality, additivity, paraphrase invariance, compression ratio, sufficiency-necessity (causal-mediation ablations).
Without these the framing is rhetoric. Spec 04 is the dedicated test.
Feed residue input → model generates 200–400 tokens. Offset-aware probe (Zhang & Zhong) measures decay at offsets 0, 50, 100, 200, 400.
Tier-A residue is hypothesized to persist longer than Tier-C. Spec 02 pre-flight tests this before Spec 01.
Per text: probe score → reference distribution → LR = P(score|H1) / P(score|H2). Multi-author: Bayesian network with writer-level random effects under event-level fixed effects.
Independent LR multiplication is wrong when writers share context.
Pre-specified propositions. Methodology details. Numerical LR + verbal equivalent. Sensitivity analyses. Named failure modes. Declared proposition level.
Without proper reporting the artifact isn't citable in any forensic-adjacent venue.
For the 2-week pilot: stages 1–5 required. Stage 6 makes the framework falsifiable. Stage 7 is novel for residue. Stage 8 drops the full Bayesian network early. Stage 9 makes it citable.
Everything above is the probing rung — infrastructure for reading a marginal, model-specific signal off activations. Two further rungs turn that infrastructure into something an agent can build with. Probing was the floor; these are the build steps.
Fine-tune a model to push a marginal residue floor toward a more robust signal — the gain only counts if it holds on held-out events and held-out controls, so it can't be a register artifact or a memorized event. If it only survives on seen events, or only against Wikipedia, it doesn't count.
Wrap the nine stages as an autonomous loop with typed contracts: collect tiered corpus → build control corpora → probe → fine-tune → self-verify → integrate into a persistent base. A new event enters; the loop runs; the base is meant to carry the capability forward.
Pre-registered bands are the loop's self-honesty mechanism. In ENFSI discipline a human pre-registers pass/fail thresholds before seeing data. In the autonomous loop that same act becomes the agent's calibration check: content-addressed bands the agent cannot move mid-run. This is the thing that stops an agent from confidently reporting a confounded model — one that passes the tier gradient but fails control-corpus isolation — as residue-capable. The bands are not yet evidence the loop works; they are what would make a "works" verdict trustworthy if one ever lands.
The methodology's name borrows from complex analysis — Cauchy's residue theorem, where local data at singularities determines global behavior. The metaphor is generative. It is also, honestly framed, partial.
A holomorphic function is rigid — knowing it at a few points often determines it everywhere. Where it has singularities, each is characterized by one number: its residue. Cauchy's theorem:
Knowing the singularities and their residues is enough to compute the integral around the whole loop.
The bet: downstream behavior Φ(x) — a probe output, a steering effect — may be similarly determined by a small number of localized "singular spans." The residue at each span is a low-dim summary of its contribution:
The residual stream is literally a sum of per-layer per-head contributions (Elhage 2021). Direct logit attribution (nostalgebraist 2020) — final residual stream is a linear sum projected onto next-token logits. Additivity exact by construction.
For arbitrary downstream Φ — a deep probe, a behavioral classification — additivity is approximate. No Cauchy theorem to invoke. Empirical near-linearity has to be tested. Falsifiable approximation, not proven identity.
The math residue language is a metaphor with a partial rigorous core: direct-effect attribution with sparse local features, motivated by complex analysis but not proven by it. The five criteria make it falsifiable.
Spec 04 is the dedicated test. Aggregate verdict gates bible decision 5: 4–5/5 → locked. 3/5 → partial revision. ≤2/5 → drop decision 5. The willingness to drop is the entire point of having criteria.
Everything above presupposes a substrate: that emotion concepts actually live as linear directions in the residual streams of contemporary LLMs. Before building forensic methodology on top, the substrate had to be confirmed independently. Five open-weight models, one common pipeline, one unexpected finding.
Sofroniew, Kauvar, Saunders et al. — "Emotion Concepts and their Function in a Large Language Model," Transformer Circuits Thread, April 2026. Emotion concepts live as linear directions in Claude Sonnet 4.5's residual stream — directions that reproduce the human affective circumplex (valence × arousal) and causally influence behavior. They published the recipe, not the data.
Same recipe, five open-weight LLMs from three labs. Same 20-emotion / 30-topic prompt template, same PCA denoising, same activation averaging. Plus statistical rigor Anthropic didn't formalize (bootstrap CIs, permutation tests, probe accuracy as a comparable scalar).
| Model | Layer | Probe acc | PC1 valence sep | Implicit top-3 | Geometric profile |
|---|---|---|---|---|---|
| Qwen2.5-1.5B-Instruct | 17/28 | 89.7% | 7.30 | 20% | Profile A · valence |
| Qwen2.5-7B-Instruct | 18/28 | 91.8% | 12.06 | 60% | Profile A · valence |
| Qwen3-8B | 23/36 | 91.0% | 29.19 | 40% | Profile A · valence |
| Llama-3.1-8B-Instruct | 16/32 | 92.1% | 2.71 | 60% | Profile B · distributed |
| Mistral-7B-v0.3 | 16/32 | 91.6% | 1.57 | 50% | Profile B · distributed |
All five pass at 89.7–92.1% probe accuracy on 20-way emotion classification (chance 5%). Substrate confirmed: emotion concepts do live as linear directions in contemporary open-weight LLM residual streams.
Beyond confirming the substrate, the replication surfaced something the single-model paper couldn't see: more than one valid way to encode emotion linearly.
One strong PC1 axis encodes positive vs negative — up to 29× separation in Qwen3-8B. Emotion compressed onto a valence ladder, arousal a weaker secondary axis. Sentiment-style tasks benefit.
Flat valence axis (PC1 separation under 3) but tighter within-cluster cohesion. Emotion spread across many cluster-defining directions. Llama matches the best Qwen on implicit-emotion identification — fine-grained tasks benefit.
Open artifact: emotion-vector-bench — 5 model results, 3050-stimulus corpus, 7-stage pipeline, denoised vectors. Anthropic published the recipe; we publish the data.
A research program is defined as much by what it refuses to claim as by what it asserts. These aren't aspirational limits — they are constitutive constraints. Each refusal corresponds to a known failure mode in an adjacent field.
The model does not "feel" anything. Emotion vectors are functional patterns of activation, nothing more. We use residue instead of emotion to sidestep the consciousness question.
Forensic linguistics has a developed authorship paradigm (Nini 2023). We do not claim "writer X wrote text Y." We claim residue is consistent with a role in an event — structural property, not identity.
We never report "this text has anxiety residue" as a categorical claim. Classification overstates certainty and bakes in priors not ours to set. We report LRs over pre-specified propositions.
The prosecutor's fallacy: "residue is one-in-a-million under H2, therefore H1 is 99.9999% true." Wrong — confuses P(E|H) with P(H|E), ignores priors. We report LR; the trier supplies the prior.
Cook et al.'s hierarchy is a fence, not a ladder. Source-level is trivial; offence-level (intent, causation) is out of scope. We live strictly at activity. Drift mid-report is the most common forensic-statistics error.
The math residue framing is a metaphor with a partial rigorous core (§04). We submit it to five criteria. If ≤2 pass, decision 5 is dropped. The willingness to drop is what distinguishes framework from rhetoric.
Apollo dropped AI Audit because labels couldn't be defended. We inherit it: when tier-A/C labels can't be ground-truthed, the corpus doesn't enter the analysis.
The contribution is the integration. We borrow MM probes (Marks & Tegmark), LR + follow-up probes (Apollo), the 171-emotion list (Anthropic), forensic statistics from a century of practice. Novelty: bringing them into one pipeline on one substrate — and, increasingly, an agent loop that runs them.
Spec 1.5 is the cautionary finding: a probe can satisfy the bible's monotonic A>B>C>D tier ordering and still be a register detector, not a residue detector. The load-bearing validation is the control-corpus isolation test — separating event-residue from ordinary and emotional first-person text. The tier gradient is demoted to a sanity check.
For the autonomous loop, ENFSI pre-registration becomes the agent's calibration check: content-addressed pass/fail bands the agent cannot move mid-run. This is what keeps the loop from confidently reporting a confounded model as residue-capable. The honest state remains a marginal, model-specific probing floor — the bands govern how a stronger claim would have to be earned, not whether one exists yet.
Naming the refusals up front means they don't get rediscovered the hard way. The positive program is more honest because the negative space is explicit.
The bet (capability vs felt-sense), active experiments, build order, the four-step walkthrough of what residue is, the two-core finance application, locked decisions, and named gaps. Compact version of the program.
Citations and inheritances mapped: Apollo, Cook et al., Carrier & Spafford, Vieweg, Marks & Tegmark, Zhang & Zhong, Anthropic emotion concepts. Where each piece comes from and what we owe each source.