Written 2026-05, before the probing programme closed. For experimental outcomes see /results.

Apollo. Forensic statistics. Digital forensics. Crisis informatics.

Each pillar contributes a different epistemological organ. None invented here. All load-bearing. The novel contribution is the integration — strip out any one and the structure fails in a specific, predictable way.

Pillar I
Apollo · linear probes
Goldowsky-Dill, Chughtai, Heimersheim, Hobbhahn · arXiv 2502.03407 · 2025
The clean template for activation-probe interpretability. Contrast pairs, L2-regularized probes, layer sweeps, AUROC + recall@1%FPR on a clean control corpus, and the willingness to drop dirty datasets.
WE BORROW: the full pipeline plus three named failure modes (spurious correlation, aggregation, mysterious).
Strip out → no mechanism for reading residue from activations.
Pillar II
Forensic statistics · LR framework
Aitken, Taroni & Bozza · ENFSI 2015 · Cook et al. 1998 · Nini 2023
A century of forensic science: quantify weight of evidence, never posterior probability of guilt. Scientist computes LR over pre-specified propositions; trier supplies the prior. Nini 2023 brought it to forensic linguistics; we bring it to residue.
WE BORROW: LR not classification, pre-specified propositions, the hierarchy (§02), ENFSI's four pillars, four named fallacies.
Strip out → claims become un-peer-reviewable in forensic-adjacent venues.
Pillar III
Digital forensics · event reconstruction
Carrier & Spafford 2003 · Casey · DFRWS · Locard 1920s
Carrier & Spafford's move: treat a digital system as a crime scene investigable by physical-forensics methods. Events are state transitions with cause, time, location, effect. Locard's principle: every contact leaves traces.
WE BORROW: event reconstruction over writer-state inference, primary vs secondary scenes (→ tier-A/C), Casey's C-scale, anti-forensic robustness (sanitization-resistant residue lives in other people's records).
Strip out → the project collapses to mind-reading the activations don't license.
Pillar IV
Crisis informatics · collective text
Vieweg, Hughes, Starbird, Palen · CHI 2010 · Palen & Liu 2007 · Olteanu 2014
15+ years of CHI/CSCW work: aggregate social-media text encodes the spatiotemporal contour of events even when individual messages are noisy. Affect shifts mark phase transitions. On-the-ground voices carry information external commentary doesn't.
WE BORROW: the precedent that collective-text → event reconstruction works, the on-the-ground / external distinction (ancestor of tier-A/C), CrisisLex-style corpus construction.
Strip out → no precedent the substrate carries signal.

Missing on purpose: affective computing, sentiment analysis, clinical psychometrics, authorship attribution. Adjacent, not lineage. Their epistemology doesn't survive our constraints.

Cook et al. 1998. We live strictly at activity level.

A 1998 paper in Science & Justice formalized something forensic scientists had been doing for decades: declare the level of claim, and don't drift mid-report. Three levels — source, activity, offence. Staying at one level is what makes the framework computable.

level
what it claims
our setting
Source · I
"This trace came from this source." Pure trace-to-source matching. The DNA at the scene matches suspect X.
Trivial. "This activation pattern was produced by reading this text" — yes, by construction. There's no inference problem here.
Activity · II
› our target
"This trace deposited during the event of interest." Requires transfer, persistence, recovery dynamics. The suspect struck the victim (without claiming intent).
Our target. "Was the writer in role R during event E?" Roles: on-the-ground participant, witness, secondhand-recipient, institutional spokesperson, retrospective observer.
Offence · III
"The suspect intended the act." Adds causation and legal/intent elements. Explicitly the trier's job, not the scientist's.
Out of scope. We never claim the writer caused the event or intended a state of mind. The residue doesn't license offence-level claims.
Worked example · yen-carry unwind · Aug 1–5 2024

Source · "Activation X was produced by reading trader Y's text." True by definition. Uninteresting.

Activity › "Residue patterns across 100 writers are ~10⁴ times more likely under [some were on-the-ground carry-trade participants] than under [all were external observers]." This is what residue can tell us. Reportable as an LR.

Offence ✕ "These writers caused the yen unwind." Out of scope. The residue doesn't say this even when the LR is extreme.

The discipline forces honesty about what residue can tell us. It can identify who was in an event. It cannot identify who caused it. That is plenty — genuinely useful and currently un-served — but it is the only thing the activations license. Drifting to offence-level is the prosecutor's fallacy in academic clothing.

How residue actually gets made.

The four-step walkthrough on the landing page is the public sketch. This is the working pipeline. Skip a stage and something measurable breaks.

The pipeline is becoming an agent loop. These nine stages were first written as a process a human researcher runs by hand. The thing we are actually building is the same sequence redrawn as agent-executable steps with typed input/output contracts — collect a tiered corpus → build control corpora → probe → fine-tune → self-verify → integrate into a persistent base. The lab's load-bearing contribution is not a single probe or model but a runnable methodology an autonomous agent executes to attempt felt-sense capability for a new event. Each event is meant to leave the system slightly better at the next. Whether that loop produces signal above the probing floor is the open question, not a settled result.

01
Define the event + propositions

Specify event scope. Pre-specify H1 and H2 at activity level. Declare proposition level. ENFSI requires pre-registration.

Without this, no LR is computable and the framework collapses to post-hoc classification.

output: proposition spec · event identifier · level declaration
02
Corpus collection

Two corpora. Event tiered A–D (A first-person on-the-ground; C institutional). Reference: the "irrelevant population" base rate — Wikipedia, recipes, code docs.

Forensic LRs need a reference population. Otherwise "more likely than what" has no answer.

output: tiered event corpus · reference corpus · methodology doc
03
Activation extraction

Forward-pass each text. Capture residual-stream activations across layers — sweep at 25/50/75% depth minimum. Normalize zero-mean unit-variance.

Apollo 27%; Zhang & Zhong 50–75% for emotion. Assuming mid-late is a known failure mode.

output: per-text per-layer normalized activation tensors
04
Probe construction

Contrast pairs (tier-A / reference, or tier-A / tier-C). Mass-mean primary (Marks & Tegmark 2024), L2 logistic-regression secondary baseline, follow-up-question probes per Apollo.

MM-primary is the May 2026 bible update. LR tilts off the actual concept direction.

output: probe directions per layer per contrast
05
Probe evaluation · control-corpus isolation

AUROC + recall@1%FPR (deployment metric). Span-level scoring. Reliability-diagram calibration. Spurious-correlation ablations vs adjacent concepts (anger / moral indignation, grief / solemnity). Load-bearing check: does the probe separate genuine event-residue from ordinary and emotional first-person control text — not just from Wikipedia?

Apollo's failure mode #1 was tracking morality not honesty — we plan for the analog. Spec 1.5 found a probe can satisfy the bible's monotonic A>B>C>D tier ordering while still being a register detector, not a residue detector. The tier gradient is now a sanity check; the control-corpus isolation test is what the claim rests on.

output: probe scorecard · calibration · ablations · control-corpus separation
06
Math residue criteria

For each probe + Φ, compute the five criteria from §04: locality, additivity, paraphrase invariance, compression ratio, sufficiency-necessity (causal-mediation ablations).

Without these the framing is rhetoric. Spec 04 is the dedicated test.

output: five criterion scores · aggregate verdict
07
Persistence-decay measurement

Feed residue input → model generates 200–400 tokens. Offset-aware probe (Zhang & Zhong) measures decay at offsets 0, 50, 100, 200, 400.

Tier-A residue is hypothesized to persist longer than Tier-C. Spec 02 pre-flight tests this before Spec 01.

output: per-tier decay curves · signed asymmetry test
08
LR computation + Bayesian aggregation

Per text: probe score → reference distribution → LR = P(score|H1) / P(score|H2). Multi-author: Bayesian network with writer-level random effects under event-level fixed effects.

Independent LR multiplication is wrong when writers share context.

output: per-text LR · event aggregate · verbal-scale equivalent
09
ENFSI-style report

Pre-specified propositions. Methodology details. Numerical LR + verbal equivalent. Sensitivity analyses. Named failure modes. Declared proposition level.

Without proper reporting the artifact isn't citable in any forensic-adjacent venue.

output: reproducible ENFSI-style technical report · the publishable artifact

For the 2-week pilot: stages 1–5 required. Stage 6 makes the framework falsifiable. Stage 7 is novel for residue. Stage 8 drops the full Bayesian network early. Stage 9 makes it citable.

03.1 Two rungs beyond probing

Everything above is the probing rung — infrastructure for reading a marginal, model-specific signal off activations. Two further rungs turn that infrastructure into something an agent can build with. Probing was the floor; these are the build steps.

› rung · amplify

Fine-tune a model to push a marginal residue floor toward a more robust signal — the gain only counts if it holds on held-out events and held-out controls, so it can't be a register artifact or a memorized event. If it only survives on seen events, or only against Wikipedia, it doesn't count.

› rung · automate

Wrap the nine stages as an autonomous loop with typed contracts: collect tiered corpus → build control corpora → probe → fine-tune → self-verify → integrate into a persistent base. A new event enters; the loop runs; the base is meant to carry the capability forward.

Pre-registered bands are the loop's self-honesty mechanism. In ENFSI discipline a human pre-registers pass/fail thresholds before seeing data. In the autonomous loop that same act becomes the agent's calibration check: content-addressed bands the agent cannot move mid-run. This is the thing that stops an agent from confidently reporting a confounded model — one that passes the tier gradient but fails control-corpus isolation — as residue-capable. The bands are not yet evidence the loop works; they are what would make a "works" verdict trustworthy if one ever lands.

The Cauchy analogy. Metaphor, not theorem.

The methodology's name borrows from complex analysis — Cauchy's residue theorem, where local data at singularities determines global behavior. The metaphor is generative. It is also, honestly framed, partial.

04.1 The actual theorem, in plain language

A holomorphic function is rigid — knowing it at a few points often determines it everywhere. Where it has singularities, each is characterized by one number: its residue. Cauchy's theorem:

∮γ f(z) dz = 2πi · Σ Res(f, ak)
Local data at singularities determines global behavior around any closed loop

Knowing the singularities and their residues is enough to compute the integral around the whole loop.

04.2 Why we borrowed it

The bet: downstream behavior Φ(x) — a probe output, a steering effect — may be similarly determined by a small number of localized "singular spans." The residue at each span is a low-dim summary of its contribution:

Φ(x) ≈ Φ0 + Σs ∈ S(x) ⟨vΦ, r(s)⟩
The linear-residue hypothesis · aspirational form

04.3 Literal vs aspirational

› literal

The residual stream is literally a sum of per-layer per-head contributions (Elhage 2021). Direct logit attribution (nostalgebraist 2020) — final residual stream is a linear sum projected onto next-token logits. Additivity exact by construction.

› aspirational

For arbitrary downstream Φ — a deep probe, a behavioral classification — additivity is approximate. No Cauchy theorem to invoke. Empirical near-linearity has to be tested. Falsifiable approximation, not proven identity.

The math residue language is a metaphor with a partial rigorous core: direct-effect attribution with sparse local features, motivated by complex analysis but not proven by it. The five criteria make it falsifiable.

04.4 The five falsifiability criteria

CRIT 01
Locality
Top-5 singular spans ≥ 90% of total causal effect. If this fails, "residue at singularities" is wrong — everything matters everywhere.
CRIT 02
Additivity
R² ≥ 0.7 between Σ⟨v_Φ, r(s)⟩ and observed Φ on held-out data. Low R² → cross-singularity interactions dominate; can't sum across writers either.
CRIT 03
Paraphrase invariance
‖r(s) − r(s')‖ / ‖r(s)‖ < 0.2 for matched paraphrases preserving singular spans. Else the probe tracks surface words, not the underlying concept.
CRIT 04
Compression ratio
ρ = (|S(x)| × residue-dim) / n < 0.1 — residue at most 10% of the text. Otherwise relabeling, not compressing.
CRIT 05
Sufficiency-and-necessity
Replace non-singular tokens → Φ preserved. Ablate any singular span → Φ shifts > τ. Tested via causal mediation (Vig 2020, ROME 2022).

Spec 04 is the dedicated test. Aggregate verdict gates bible decision 5: 4–5/5 → locked. 3/5 → partial revision. ≤2/5 → drop decision 5. The willingness to drop is the entire point of having criteria.

Replicating Anthropic. Five open-weight models.

Everything above presupposes a substrate: that emotion concepts actually live as linear directions in the residual streams of contemporary LLMs. Before building forensic methodology on top, the substrate had to be confirmed independently. Five open-weight models, one common pipeline, one unexpected finding.

05.1 What Anthropic established

Sofroniew, Kauvar, Saunders et al. — "Emotion Concepts and their Function in a Large Language Model," Transformer Circuits Thread, April 2026. Emotion concepts live as linear directions in Claude Sonnet 4.5's residual stream — directions that reproduce the human affective circumplex (valence × arousal) and causally influence behavior. They published the recipe, not the data.

05.2 What we replicated

Same recipe, five open-weight LLMs from three labs. Same 20-emotion / 30-topic prompt template, same PCA denoising, same activation averaging. Plus statistical rigor Anthropic didn't formalize (bootstrap CIs, permutation tests, probe accuracy as a comparable scalar).

Model Layer Probe acc PC1 valence sep Implicit top-3 Geometric profile
Qwen2.5-1.5B-Instruct 17/28 89.7% 7.30 20% Profile A · valence
Qwen2.5-7B-Instruct 18/28 91.8% 12.06 60% Profile A · valence
Qwen3-8B 23/36 91.0% 29.19 40% Profile A · valence
Llama-3.1-8B-Instruct 16/32 92.1% 2.71 60% Profile B · distributed
Mistral-7B-v0.3 16/32 91.6% 1.57 50% Profile B · distributed

All five pass at 89.7–92.1% probe accuracy on 20-way emotion classification (chance 5%). Substrate confirmed: emotion concepts do live as linear directions in contemporary open-weight LLM residual streams.

05.3 The two geometric profiles

Beyond confirming the substrate, the replication surfaced something the single-model paper couldn't see: more than one valid way to encode emotion linearly.

Profile A Dominant valence axis

Qwen family · 2.5-1.5B, 2.5-7B, 3-8B

One strong PC1 axis encodes positive vs negative — up to 29× separation in Qwen3-8B. Emotion compressed onto a valence ladder, arousal a weaker secondary axis. Sentiment-style tasks benefit.

Profile B Distributed clusters

Llama-3.1-8B · Mistral-7B-v0.3

Flat valence axis (PC1 separation under 3) but tighter within-cluster cohesion. Emotion spread across many cluster-defining directions. Llama matches the best Qwen on implicit-emotion identification — fine-grained tasks benefit.

05.4 Why this matters for the residue work

Open artifact: emotion-vector-bench — 5 model results, 3050-stimulus corpus, 7-stage pipeline, denoised vectors. Anthropic published the recipe; we publish the data.

The negative space is a feature.

A research program is defined as much by what it refuses to claim as by what it asserts. These aren't aspirational limits — they are constitutive constraints. Each refusal corresponds to a known failure mode in an adjacent field.

DISC 01
not consciousness inference

The model does not "feel" anything. Emotion vectors are functional patterns of activation, nothing more. We use residue instead of emotion to sidestep the consciousness question.

DISC 02
not authorship attribution

Forensic linguistics has a developed authorship paradigm (Nini 2023). We do not claim "writer X wrote text Y." We claim residue is consistent with a role in an event — structural property, not identity.

DISC 03
not classification labels

We never report "this text has anxiety residue" as a categorical claim. Classification overstates certainty and bakes in priors not ours to set. We report LRs over pre-specified propositions.

DISC 04
not posterior probability of causation

The prosecutor's fallacy: "residue is one-in-a-million under H2, therefore H1 is 99.9999% true." Wrong — confuses P(E|H) with P(H|E), ignores priors. We report LR; the trier supplies the prior.

DISC 05
activity-level only

Cook et al.'s hierarchy is a fence, not a ladder. Source-level is trivial; offence-level (intent, causation) is out of scope. We live strictly at activity. Drift mid-report is the most common forensic-statistics error.

DISC 06
metaphor, not theorem

The math residue framing is a metaphor with a partial rigorous core (§04). We submit it to five criteria. If ≤2 pass, decision 5 is dropped. The willingness to drop is what distinguishes framework from rhetoric.

DISC 07
drop dirty datasets

Apollo dropped AI Audit because labels couldn't be defended. We inherit it: when tier-A/C labels can't be ground-truthed, the corpus doesn't enter the analysis.

DISC 08
not a novel probe or architecture

The contribution is the integration. We borrow MM probes (Marks & Tegmark), LR + follow-up probes (Apollo), the 171-emotion list (Anthropic), forensic statistics from a century of practice. Novelty: bringing them into one pipeline on one substrate — and, increasingly, an agent loop that runs them.

DISC 09
tier gradient is not sufficient

Spec 1.5 is the cautionary finding: a probe can satisfy the bible's monotonic A>B>C>D tier ordering and still be a register detector, not a residue detector. The load-bearing validation is the control-corpus isolation test — separating event-residue from ordinary and emotional first-person text. The tier gradient is demoted to a sanity check.

DISC 10
bands locked before the agent sees data

For the autonomous loop, ENFSI pre-registration becomes the agent's calibration check: content-addressed pass/fail bands the agent cannot move mid-run. This is what keeps the loop from confidently reporting a confounded model as residue-capable. The honest state remains a marginal, model-specific probing floor — the bands govern how a stronger claim would have to be earned, not whether one exists yet.

Naming the refusals up front means they don't get rediscovered the hard way. The positive program is more honest because the negative space is explicit.