Experimental record · complete · updated 2026.08.17

Results

Every experiment run under the programme, the criteria locked before execution, and what came back. Including the four that stopped early, the two that were retracted, and the one number that has to be corrected in public.

What this page is for. Across roughly twenty-five pre-registered numeric criteria in ten executed experiments, the count that passed and still stands is small. Not one criterion was quietly relaxed to manufacture a pass, and five hypotheses were amended after seeing data — four of them toward stricter. That record is the deliverable here, not a finding.

Every verdict below is quoted from its own results file, and the file is named. Two outcomes have no committed results artifact and are labelled as such rather than stated as results.

Phase 1 · probingDoes residue exist as a readable signal?

Spec 01 · 2026-05Tier-A vs Tier-C residue transfer 0/2 literalsupplementary pass
Locked criteria
(1) Mass-mean probe degrades ≤0.05 while logistic degrades >0.10 across events. (2) Tier-A probes transfer better than Tier-C by ≥0.05 AUROC.
Outcome
Neither met. Both probes saturated on the locked contrast — event text against Wikipedia, recipes and code documentation is trivially separable, so the asymmetric pattern the criteria looked for had no room to appear. Prediction #2 gap at best layer: +0.008 against a required +0.05.
Supplementary
A harder contrast (train tier-A vs tier-C within Event A, transfer to Event B) was run with bands locked in the script before first execution. Consistent PASS across all 12 layer×probe settings on Qwen3-8B and all 9 on Llama-3.1-8B. Best variant both families, mass-mean unwhitened: 0.980 in-distribution / 0.980 transfer — zero degradation.
Complication
The tier gradient came out A > D > B > C, not the predicted A > B > C > D. That ordering suggests the probe is substantially a personal-versus-institutional register detector. Spec 1.5 was built to settle it.
Verdict
“0/2 by literal spec, consistent supplementary PASS across two model families. The signal exists; the spec's locked contrast couldn't see it.”
synthesis/results/01-tier-a-vs-tier-c-results.md
Spec 02 · 2026-05-27Persistence-decay on yen-carry fail
Locked rule
At offset ≥100 generated tokens, tier-A probe accuracy exceeds tier-C by ≥0.10, Bonferroni-significant.
Outcome
Not met under either of two runs. Run one (enable_thinking=False, a divergence from spec §4) landed closest to Fail Mode B — tier-C persisted longer, 96% against 33% at offsets ≥150, the opposite of the hypothesis. The spec-compliant re-run (enable_thinking=True) landed closer to partial Fail Mode A — curves overlap, 83% against 90%, gaps noisy at ±0.20 and inconsistent in direction.
Cross-run finding
“Persistence-into-continuation is dominated by LLM behavior, not by intrinsic residue persistence.” At offset 0, detection is robust under both settings (≥80% per tier). Downstream offsets measure a mixture of residue and conversational mode that cannot be separated in this setup.
Action
Spec 01 prediction #3 dropped 2026-05-27. The amendment holds under both runs.
synthesis/results/02-persistence-decay-results.md
Spec 1.5 · 2026-05-28Register / valence control — the decisive one inconclusive
Question
Was Spec 01's clean supplementary separation reading residue, or just personal-versus-institutional voice? The easy institutional foil is swapped for a hard one: ordinary personal writing, same register, no event behind it.
Results
Qwen3-8B L14 — 0.791 / 0.767 — pass
Qwen2.5-7B L7 — 0.587 / 0.416 — fail
Llama-3.1-8B L8/L16/L24 — 0.43–0.57 / 0.34–0.50 — fail at all three layers, worse with depth
Verdict
Spec §2 requires agreement across families to count. “EXECUTED — CROSS-FAMILY DISAGREE → INCONCLUSIVE per spec §2.”
What it retracted
“The Spec 01 'cross-family confirmation' claim is retracted.” Both families had passed the supplementary contrast because that contrast could not distinguish residue from register. The purpose-built control corpora broke the symmetry and the families diverged sharply.
What survives
Substrate availability is a measurable per-model property, not a property of open-weight models in general. Qwen3-8B's lower CI bound sits near chance, so its pass is fragile.
synthesis/results/01.5-register-control-results.md

Phase 2 · mission trackCan the floor be raised, and can the loop run itself?

Spec CT “Chemtech” · 2026-05-30Residue structure + designed-vs-genuine CT-A mixedCT-B pass
CT-A structure
MIXED. Top-5 tokens carry 58% of contribution mass, but removing them retains 99.7% of full-text AUROC (0.980 → 0.977). The signal is redundantly distributed. Residue concentrates mildly in the back third of text (0.38 vs 0.33 front, 0.30 middle).
CT-A scope correction
Same-day audit: the verdict is accurate but its label was too generous. It measured the tier-A-vs-tier-C contrast — the register-laden one. The honest claim is “the register-laden contrast is redundantly distributed,” not that pure residue is. A follow-up on the residue contrast could not be run cleanly: a dedicated residue probe saturates at 1.0, so redundancy is uninformative there.
CT-B baseline test
PASS on Qwen3-8B. Residue probe 0.980 transfer AUROC against a TF-IDF + length + Flesch + type-token-ratio + function-word surface baseline at 0.860. Margin +0.120 against a locked threshold of 0.10; permutation p = 0.005 against a locked 0.01. Both thresholds fixed in advance, both cleared.
Cross-family
MIDDLE. Qwen2.5-7B and Llama both clear the margin (+0.112, +0.120) but miss the strict p-cutoff (0.015, 0.012).
Caveat carried
“CT-B's PASS means 'residue probe beats surface features,' not 'residue probe is pure residue.'” It inherits Spec 1.5's register confound.
synthesis/results/CT-chemtech-results.md
Spec M1 · 2026-05-30Fine-tune Qwen3-8B — raise the floor bands passambiguous
By locked bands
PASS. AUROC 1.000 against tier-A-vs-calm and 1.000 against tier-A-vs-emotional, both saturating the locked thresholds of ≥0.90 and ≥0.85. Pipeline ran end to end: LoRA trained 30 minutes on MPS, splits hashed before train and eval, no leakage.
By the discriminating contrast
Lift +0.004. On held-out within-Event-D tier-A vs tier-C, base Qwen3-8B already scores 0.966 [CI 0.895, 1.000]; the fine-tuned model scores 0.970. The locked PASS was the base model's existing geometry pushed to ceiling on an easy contrast.
Forgetting
DEGRADED. Mean perplexity ratio 1.79; one probe blew up 3.85×. Not catastrophic.
Decision
“Do not proceed to M2 with this LoRA.”
Revised verdict
A same-day audit revised this to AMBIGUOUS: “Both this memo's 'NULL / base already had it' conclusion and the audit's first counter-claim ('a clean win') were over-confident in opposite directions.” The decision did not change; the reason sharpened.
synthesis/results/M1-residue-finetune-results.md · revised by probe-isolation-limit-2026-05-30.md §3.2
Spec M2 “shorekeeper”Automate the loop — the stated deliverable never started
Status
Never started, and correctly so. M2 was gated on M1 passing its discriminating contrast. M1 returned +0.004 and instructed against proceeding. The gate never opened.
Note
This was the programme's headline deliverable — an agent-driven loop that collects a corpus, builds controls, probes, fine-tunes, and self-verifies against locked bands. Running it on a LoRA that showed no capability shift would have been the violation, not the achievement.
synthesis/specs/M2-autonomous-loop/ — specification only, no results file

The characterizationWhere the probing programme actually ended

Audit · 2026-05-30Probe isolation limit completed negative

Six independent attempts to isolate residue from register and event-presence with linear probes hit the same wall. The audit that closed the phase states the result plainly:

“Six attempts, same wall. That's not failure — it's a characterization: residue does not exist as a clean separable linear feature in these representations. It is distributed and entangled with event-presence and register.”

“The instrument can't resolve it, and the program should change instruments.”

This is the one question in the programme carried to a conclusion. It is a negative result with a stated mechanism and a decision attached, and it is the reason the programme moved from probing to reconstruction.

synthesis/results/probe-isolation-limit-2026-05-30.md

R1 · 2026-06Reconstruction — a different instrument

R1 Phase 1 · syntheticReconstruction beats a facts baseline — on synthetic worlds gates pass
Rigor gates
All PASS. Facts baseline affect-neutral and fact-complete; reconstruction injects no facts absent from residue; strict causal and temporal masking verified across three axes; β=0 negative control clean with oracle ceilings computed first.
Wins
Three of six cells beat the facts baseline with intervals excluding zero:
axistargetdelta95% CI
longitudinalresidue+0.171[+0.133, +0.209]
longitudinalaction+0.040[+0.015, +0.067]
cross-sectionalaction+0.028[+0.007, +0.049]
Stated limit
The report's own caveat: the method “only wins where the structure is” present. Cross-sectional action beats facts only with genuine collective structure in the world.
Scope
Synthetic worlds with a known answer key. No real-data verdict line was locked and none was run at this stage, by design.
residue-lab/PHASE1-REPORT.md · results/phase1/rigor_gate.txt
R1 Phase 1.5 · 2026-06-02Put a real local model in the scoring loop gate failed twice
Outcome
The calibration gate failed with a 1.5B scorer, then failed again, harder, when re-scored with 4-bit Qwen2.5-7B. Both reconstruction and facts lost to the prior. “The judge stays VOID.”
Honest note in the report
By the strict letter of the pre-registered control it passes — “but this passing verdict is worth little given the calibration failure above.”
residue-lab/PHASE1_5-REPORT.md
R1 forensicLocked LR gates on the yen-carry corpus borderlineunfinished

Provenance warning. These figures were reported in-session and no committed results file exists on disk — only the gate machinery (gates.py, gate2.py, calibrate.py). Treat as reported, not as an artifact, until re-run.

Gate 1
G1a passes (Cllr mean 0.55, meets ≤0.80 in 90% of splits). G1c robust (directional purity H1≈0.86 / H2≈0.96). G1b sits on the bar — ECE mean 0.103 against a ≤0.10 threshold, met in only about half of splits. Full gate passes 5 of 10 seeds; a 40-seed sweep confirms roughly 60%.
Amendment
The verbal scale was downgraded from “strong” (LR≥100) to “moderate” (LR≥10) on grounds of data scale (N≈99) — a pre-registered threshold moved after seeing the number. The amendment retracts its own framing: “This is a downgrade, not a relocation; the earlier wording overstated it.”
The decisive test
Never run. “Does R1 actually pass on AWS? This is the question... The reconstruction pipeline itself hasn't been run on the post-cutoff AWS corpus.” The held-out event remains untouched.
reported in raw session transcripts, 2026-05-30 to 2026-06-01 · no results file committed

R2 · 2026-08Acquisition methodology as a training variable

R2 · 2026-08-05Does how you collect a corpus change what a model learns? retracted twice
Pre-registration
Signed 2026-08-05 before any training run, sha256 beginning 5d2f1721d119686f. Three conditions, all required to hold.
#condition, fixed in advanceresultverdict
1event-bootstrap 95% CI excludes 0[+0.0224, +0.0740]PASS
2point estimate ≥ 5.0 pp4.75 ppFAIL
3effect exceeds across-seed spread4.75 vs 3.34 / 2.46PASS
Original verdict
SUGGESTIVE, NOT ESTABLISHED. Two of three. The effect fell 0.25 pp short. The results file states why the rule existed: “read after the fact, 4.75 pp is trivially arguable into a win.”
Retraction 1 · 08-06
Grading artifact. The scorer accepted a bare year as a full date, and one arm's corpus taught models to emit years. Year-emission rates: directed arm 43.9 / 56.0 / 54.6% against undirected 21.8 / 24.2 / 24.9%. With date-valued golds removed, the estimate becomes −0.58 pp, CI [−3.67, +1.93] — indistinguishable from zero. The defect is twelve lines in pipeline/r2/grade.py.
Retraction 2 · 08-08
Independent fidelity confound. Only 44.6% of primary-endpoint facts survived into both training corpora, giving one arm a +11.4 pp structural head start before any learning. The design could not have measured what it claimed to.
Replication
Cross-family replication did not replicate.
Standing
Sections 1–9 are preserved unaltered as the original record and should be read as retracted. No R2 number should be resurrected.
residue-lab/R2/RESULTS.md §1, §10, §11 · R2/PREREGISTRATION.lock

Two open defects in the R2 record, stated rather than smoothed.

§10.6 cites R2/AUDIT.md. That file does not exist on disk. The citation needs removing or repointing.

R2/results/REPORT.txt reports events=49 mean diff=+0.041 CI=[+0.007,+0.073] while RESULTS.md §2 reports +0.0475 CI [+0.0224,+0.0740]. These are two different estimators — a per-event mean over 49 events against a pooled figure — which is benign but undocumented, and reads as an inconsistency to anyone checking.

R3 · 2026-08-09Arrangement — instrument first, experiment later

R3 evalSealed-corpus evaluation set built · 0 of 5 gates

After R2 was retracted for a measurement defect, the next step was to rebuild the instrument before running anything through it. Fifty independent agents each deep-read one event's primary documents under a sealed-corpus rule — no web access, no outside knowledge — wrote an understanding field first, then authored items with verbatim evidence spans and named source files.

metricvaluecheck
event files50counted
total items1,265counted
multi-document780 (61.7%)counted
single-document485 (38.3%)counted
understanding written first50 / 50counted
items missing evidence0counted
quant_compare224counted
temporal_order153counted
entity_link149counted
contradiction144counted
causal_chain110counted
Status
Zero of five validation gates run. Base-model filter, fidelity canary, blind verification, freeze-and-hash, then arms and training — none executed. This is an asset, not a result, and it is not presented as one.
Known blocker
The training experiment it was built for is not viable on the available hardware: an 8,192-token context ceiling means none of the fifty events fit whole.
residue-lab/R3/eval/authored/E01..E50.json · counts reproduced directly from the JSON, 2026-08-17

Outside the residue programmeThe one result that replicates

emotion-vector-bench · 2026-05Emotion geometry is model-invariant replicates

Anthropic's April 2026 work showed emotion concepts live as linear directions in the residual stream, on closed Sonnet 4.5. This reproduces the methodology on five open-weight models across three labs and a five-fold parameter range, in one command.

modellab20-way probe× chance
Llama-3.1-8B-InstructMeta91.5% ±0.918.3×
Mistral-7B-Instruct-v0.3Mistral91.3% ±0.818.3×
Qwen2.5-7B-InstructAlibaba90.5% ±1.018.1×
Qwen3-8BAlibaba90.2% ±0.618.0×
Qwen2.5-1.5B-InstructAlibaba86.9% ±1.317.4×
Finding
The whole spread across three labs and a 1.5B-to-8B range is 4.6 percentage points. Within-versus-cross-emotion cohesion spans 0.215–0.252 — a 1.18× spread. Cross-layer stability is 0.96–0.99 for all five. A 1.5B model lands within five points of an 8B.
Full detail
emotion-vector-bench →

Correction, 2026-08-17. The repository's published headline — an “18× spread in valence-axis strength” separating Qwen from Llama and Mistral into two geometric profiles — is a units artifact and is being withdrawn.

The PC1 separation figure was computed on unnormalized vectors. Its correlation with each model's mean vector L2 norm is r = 0.9896 (R² = 0.979). Normalized, 18.58× becomes 1.40×. Qwen3-8B's mean L2 norm is 17.4; Mistral's is 1.1 — that is the entire effect.

The probe accuracies above are computed independently and are unaffected. The corrected reading — near-identical geometry across labs — is a stronger claim than the one being withdrawn.

emotion-vector-bench/results/_comparison.json · probe_results.json ×5 · recomputed from raw_vectors.npz, 2026-08-17

SummaryThe whole record on one screen

experimentlocked criteriaoutcomewhy it stopped
Spec 012 predictions0/2 literal; supplementary PASS 21/21 settingscontrast too easy; probes saturated
Spec 021 hypothesisnot met, both runs; direction reversed in oneLLM mode dominates the signal
Spec 1.5cross-family agreementINCONCLUSIVE; retracted Spec 01's cross-family claimfamilies disagreed
Spec 031 thresholdnever executedcorpus never cleared
Spec 045 criteriano committed results fileparked
Spec CT2 thresholdsCT-A MIXED (label corrected); CT-B PASS, margin +0.120, p=0.005completed
Spec M12 bandsbands PASS at 1.000; lift +0.004; AMBIGUOUSno headroom in the test
Spec M2gated on M1never startedgate never opened
Probe isolation—characterization: residue is not a clean linear featuresix attempts, one wall
R1 Phase 1rigor gates + 6 cellsall gates PASS; 3 of 6 cells beat factscompleted as scoped
R1 Phase 1.5calibration gatefailed twice; judge VOIDscorer not calibratable
R1 forensic3 gatesG1b on the bar, 5/10 seeds; reported, no artifactheld-out event never run
R23 conditions2/3 → SUGGESTIVE; then retracted twicegrader defect, then fidelity confound
R35 gateseval built, 1,265 items; 0/5 gateshardware ceiling
emotion-vector-benchreplicationreplicates on 5 models; headline correctedcompleted

Read the fourth column. In four cases a pre-registered rule said stop and it was obeyed. In one case the instrument was characterized and the programme changed instruments. In two cases the work is genuinely unfinished, and both are named above rather than described as parked.