Experimental record · complete · updated 2026.08.17
Results
Every experiment run under the programme, the criteria locked before execution, and what came back. Including the four that stopped early, the two that were retracted, and the one number that has to be corrected in public.
What this page is for. Across roughly twenty-five pre-registered numeric criteria in ten executed experiments, the count that passed and still stands is small. Not one criterion was quietly relaxed to manufacture a pass, and five hypotheses were amended after seeing data — four of them toward stricter. That record is the deliverable here, not a finding.
Every verdict below is quoted from its own results file, and the file is named. Two outcomes have no committed results artifact and are labelled as such rather than stated as results.
Phase 1 · probingDoes residue exist as a readable signal?
Spec 01 · 2026-05Tier-A vs Tier-C residue transfer0/2 literalsupplementary pass
Locked criteria
(1) Mass-mean probe degrades ≤0.05 while logistic degrades >0.10 across events. (2) Tier-A probes transfer better than Tier-C by ≥0.05 AUROC.
Outcome
Neither met. Both probes saturated on the locked contrast — event text against Wikipedia, recipes and code documentation is trivially separable, so the asymmetric pattern the criteria looked for had no room to appear. Prediction #2 gap at best layer: +0.008 against a required +0.05.
Supplementary
A harder contrast (train tier-A vs tier-C within Event A, transfer to Event B) was run with bands locked in the script before first execution. Consistent PASS across all 12 layer×probe settings on Qwen3-8B and all 9 on Llama-3.1-8B. Best variant both families, mass-mean unwhitened: 0.980 in-distribution / 0.980 transfer — zero degradation.
Complication
The tier gradient came out A > D > B > C, not the predicted A > B > C > D. That ordering suggests the probe is substantially a personal-versus-institutional register detector. Spec 1.5 was built to settle it.
Verdict
“0/2 by literal spec, consistent supplementary PASS across two model families. The signal exists; the spec's locked contrast couldn't see it.”
synthesis/results/01-tier-a-vs-tier-c-results.md
Spec 02 · 2026-05-27Persistence-decay on yen-carryfail
Locked rule
At offset ≥100 generated tokens, tier-A probe accuracy exceeds tier-C by ≥0.10, Bonferroni-significant.
Outcome
Not met under either of two runs. Run one (enable_thinking=False, a divergence from spec §4) landed closest to Fail Mode B — tier-C persisted longer, 96% against 33% at offsets ≥150, the opposite of the hypothesis. The spec-compliant re-run (enable_thinking=True) landed closer to partial Fail Mode A — curves overlap, 83% against 90%, gaps noisy at ±0.20 and inconsistent in direction.
Cross-run finding
“Persistence-into-continuation is dominated by LLM behavior, not by intrinsic residue persistence.” At offset 0, detection is robust under both settings (≥80% per tier). Downstream offsets measure a mixture of residue and conversational mode that cannot be separated in this setup.
Action
Spec 01 prediction #3 dropped 2026-05-27. The amendment holds under both runs.
synthesis/results/02-persistence-decay-results.md
Spec 1.5 · 2026-05-28Register / valence control — the decisive oneinconclusive
Question
Was Spec 01's clean supplementary separation reading residue, or just personal-versus-institutional voice? The easy institutional foil is swapped for a hard one: ordinary personal writing, same register, no event behind it.
Results
Qwen3-8B L14 — 0.791 / 0.767 — pass
Qwen2.5-7B L7 — 0.587 / 0.416 — fail
Llama-3.1-8B L8/L16/L24 — 0.43–0.57 / 0.34–0.50 — fail at all three layers, worse with depth
Verdict
Spec §2 requires agreement across families to count. “EXECUTED — CROSS-FAMILY DISAGREE → INCONCLUSIVE per spec §2.”
What it retracted
“The Spec 01 'cross-family confirmation' claim is retracted.” Both families had passed the supplementary contrast because that contrast could not distinguish residue from register. The purpose-built control corpora broke the symmetry and the families diverged sharply.
What survives
Substrate availability is a measurable per-model property, not a property of open-weight models in general. Qwen3-8B's lower CI bound sits near chance, so its pass is fragile.
MIXED. Top-5 tokens carry 58% of contribution mass, but removing them retains 99.7% of full-text AUROC (0.980 → 0.977). The signal is redundantly distributed. Residue concentrates mildly in the back third of text (0.38 vs 0.33 front, 0.30 middle).
CT-A scope correction
Same-day audit: the verdict is accurate but its label was too generous. It measured the tier-A-vs-tier-C contrast — the register-laden one. The honest claim is “the register-laden contrast is redundantly distributed,” not that pure residue is. A follow-up on the residue contrast could not be run cleanly: a dedicated residue probe saturates at 1.0, so redundancy is uninformative there.
CT-B baseline test
PASS on Qwen3-8B. Residue probe 0.980 transfer AUROC against a TF-IDF + length + Flesch + type-token-ratio + function-word surface baseline at 0.860. Margin +0.120 against a locked threshold of 0.10; permutation p = 0.005 against a locked 0.01. Both thresholds fixed in advance, both cleared.
Cross-family
MIDDLE. Qwen2.5-7B and Llama both clear the margin (+0.112, +0.120) but miss the strict p-cutoff (0.015, 0.012).
Caveat carried
“CT-B's PASS means 'residue probe beats surface features,' not 'residue probe is pure residue.'” It inherits Spec 1.5's register confound.
synthesis/results/CT-chemtech-results.md
Spec M1 · 2026-05-30Fine-tune Qwen3-8B — raise the floorbands passambiguous
By locked bands
PASS. AUROC 1.000 against tier-A-vs-calm and 1.000 against tier-A-vs-emotional, both saturating the locked thresholds of ≥0.90 and ≥0.85. Pipeline ran end to end: LoRA trained 30 minutes on MPS, splits hashed before train and eval, no leakage.
By the discriminating contrast
Lift +0.004. On held-out within-Event-D tier-A vs tier-C, base Qwen3-8B already scores 0.966 [CI 0.895, 1.000]; the fine-tuned model scores 0.970. The locked PASS was the base model's existing geometry pushed to ceiling on an easy contrast.
Forgetting
DEGRADED. Mean perplexity ratio 1.79; one probe blew up 3.85×. Not catastrophic.
Decision
“Do not proceed to M2 with this LoRA.”
Revised verdict
A same-day audit revised this to AMBIGUOUS: “Both this memo's 'NULL / base already had it' conclusion and the audit's first counter-claim ('a clean win') were over-confident in opposite directions.” The decision did not change; the reason sharpened.
synthesis/results/M1-residue-finetune-results.md · revised by probe-isolation-limit-2026-05-30.md §3.2
Spec M2 “shorekeeper”Automate the loop — the stated deliverablenever started
Status
Never started, and correctly so. M2 was gated on M1 passing its discriminating contrast. M1 returned +0.004 and instructed against proceeding. The gate never opened.
Note
This was the programme's headline deliverable — an agent-driven loop that collects a corpus, builds controls, probes, fine-tunes, and self-verifies against locked bands. Running it on a LoRA that showed no capability shift would have been the violation, not the achievement.
synthesis/specs/M2-autonomous-loop/ — specification only, no results file
The characterizationWhere the probing programme actually ended
Six independent attempts to isolate residue from register and event-presence with linear probes hit the same wall. The audit that closed the phase states the result plainly:
“Six attempts, same wall. That's not failure — it's a characterization: residue does not exist as a clean separable linear feature in these representations. It is distributed and entangled with event-presence and register.”
“The instrument can't resolve it, and the program should change instruments.”
This is the one question in the programme carried to a conclusion. It is a negative result with a stated mechanism and a decision attached, and it is the reason the programme moved from probing to reconstruction.
R1 · 2026-06Reconstruction — a different instrument
R1 Phase 1 · syntheticReconstruction beats a facts baseline — on synthetic worldsgates pass
Rigor gates
All PASS. Facts baseline affect-neutral and fact-complete; reconstruction injects no facts absent from residue; strict causal and temporal masking verified across three axes; β=0 negative control clean with oracle ceilings computed first.
Wins
Three of six cells beat the facts baseline with intervals excluding zero:
axis
target
delta
95% CI
longitudinal
residue
+0.171
[+0.133, +0.209]
longitudinal
action
+0.040
[+0.015, +0.067]
cross-sectional
action
+0.028
[+0.007, +0.049]
Stated limit
The report's own caveat: the method “only wins where the structure is” present. Cross-sectional action beats facts only with genuine collective structure in the world.
Scope
Synthetic worlds with a known answer key. No real-data verdict line was locked and none was run at this stage, by design.
R1 Phase 1.5 · 2026-06-02Put a real local model in the scoring loopgate failed twice
Outcome
The calibration gate failed with a 1.5B scorer, then failed again, harder, when re-scored with 4-bit Qwen2.5-7B. Both reconstruction and facts lost to the prior. “The judge stays VOID.”
Honest note in the report
By the strict letter of the pre-registered control it passes — “but this passing verdict is worth little given the calibration failure above.”
residue-lab/PHASE1_5-REPORT.md
R1 forensicLocked LR gates on the yen-carry corpusborderlineunfinished
Provenance warning. These figures were reported in-session and no committed results file exists on disk — only the gate machinery (gates.py, gate2.py, calibrate.py). Treat as reported, not as an artifact, until re-run.
Gate 1
G1a passes (Cllr mean 0.55, meets ≤0.80 in 90% of splits). G1c robust (directional purity H1≈0.86 / H2≈0.96). G1b sits on the bar — ECE mean 0.103 against a ≤0.10 threshold, met in only about half of splits. Full gate passes 5 of 10 seeds; a 40-seed sweep confirms roughly 60%.
Amendment
The verbal scale was downgraded from “strong” (LR≥100) to “moderate” (LR≥10) on grounds of data scale (N≈99) — a pre-registered threshold moved after seeing the number. The amendment retracts its own framing: “This is a downgrade, not a relocation; the earlier wording overstated it.”
The decisive test
Never run.“Does R1 actually pass on AWS? This is the question... The reconstruction pipeline itself hasn't been run on the post-cutoff AWS corpus.” The held-out event remains untouched.
reported in raw session transcripts, 2026-05-30 to 2026-06-01 · no results file committed
R2 · 2026-08Acquisition methodology as a training variable
R2 · 2026-08-05Does how you collect a corpus change what a model learns?retracted twice
Pre-registration
Signed 2026-08-05 before any training run, sha256 beginning 5d2f1721d119686f. Three conditions, all required to hold.
#
condition, fixed in advance
result
verdict
1
event-bootstrap 95% CI excludes 0
[+0.0224, +0.0740]
PASS
2
point estimate ≥ 5.0 pp
4.75 pp
FAIL
3
effect exceeds across-seed spread
4.75 vs 3.34 / 2.46
PASS
Original verdict
SUGGESTIVE, NOT ESTABLISHED. Two of three. The effect fell 0.25 pp short. The results file states why the rule existed: “read after the fact, 4.75 pp is trivially arguable into a win.”
Retraction 1 · 08-06
Grading artifact. The scorer accepted a bare year as a full date, and one arm's corpus taught models to emit years. Year-emission rates: directed arm 43.9 / 56.0 / 54.6% against undirected 21.8 / 24.2 / 24.9%. With date-valued golds removed, the estimate becomes −0.58 pp, CI [−3.67, +1.93] — indistinguishable from zero. The defect is twelve lines in pipeline/r2/grade.py.
Retraction 2 · 08-08
Independent fidelity confound. Only 44.6% of primary-endpoint facts survived into both training corpora, giving one arm a +11.4 pp structural head start before any learning. The design could not have measured what it claimed to.
Replication
Cross-family replication did not replicate.
Standing
Sections 1–9 are preserved unaltered as the original record and should be read as retracted. No R2 number should be resurrected.
Two open defects in the R2 record, stated rather than smoothed.
§10.6 cites R2/AUDIT.md. That file does not exist on disk. The citation needs removing or repointing.
R2/results/REPORT.txt reports events=49 mean diff=+0.041 CI=[+0.007,+0.073] while RESULTS.md §2 reports +0.0475 CI [+0.0224,+0.0740]. These are two different estimators — a per-event mean over 49 events against a pooled figure — which is benign but undocumented, and reads as an inconsistency to anyone checking.
R3 · 2026-08-09Arrangement — instrument first, experiment later
R3 evalSealed-corpus evaluation setbuilt · 0 of 5 gates
After R2 was retracted for a measurement defect, the next step was to rebuild the instrument before running anything through it. Fifty independent agents each deep-read one event's primary documents under a sealed-corpus rule — no web access, no outside knowledge — wrote an understanding field first, then authored items with verbatim evidence spans and named source files.
metric
value
check
event files
50
counted
total items
1,265
counted
multi-document
780 (61.7%)
counted
single-document
485 (38.3%)
counted
understanding written first
50 / 50
counted
items missing evidence
0
counted
quant_compare
224
counted
temporal_order
153
counted
entity_link
149
counted
contradiction
144
counted
causal_chain
110
counted
Status
Zero of five validation gates run. Base-model filter, fidelity canary, blind verification, freeze-and-hash, then arms and training — none executed. This is an asset, not a result, and it is not presented as one.
Known blocker
The training experiment it was built for is not viable on the available hardware: an 8,192-token context ceiling means none of the fifty events fit whole.
residue-lab/R3/eval/authored/E01..E50.json · counts reproduced directly from the JSON, 2026-08-17
Outside the residue programmeThe one result that replicates
emotion-vector-bench · 2026-05Emotion geometry is model-invariantreplicates
Anthropic's April 2026 work showed emotion concepts live as linear directions in the residual stream, on closed Sonnet 4.5. This reproduces the methodology on five open-weight models across three labs and a five-fold parameter range, in one command.
model
lab
20-way probe
× chance
Llama-3.1-8B-Instruct
Meta
91.5% ±0.9
18.3×
Mistral-7B-Instruct-v0.3
Mistral
91.3% ±0.8
18.3×
Qwen2.5-7B-Instruct
Alibaba
90.5% ±1.0
18.1×
Qwen3-8B
Alibaba
90.2% ±0.6
18.0×
Qwen2.5-1.5B-Instruct
Alibaba
86.9% ±1.3
17.4×
Finding
The whole spread across three labs and a 1.5B-to-8B range is 4.6 percentage points. Within-versus-cross-emotion cohesion spans 0.215–0.252 — a 1.18× spread. Cross-layer stability is 0.96–0.99 for all five. A 1.5B model lands within five points of an 8B.
Correction, 2026-08-17. The repository's published headline — an “18× spread in valence-axis strength” separating Qwen from Llama and Mistral into two geometric profiles — is a units artifact and is being withdrawn.
The PC1 separation figure was computed on unnormalized vectors. Its correlation with each model's mean vector L2 norm is r = 0.9896 (R² = 0.979). Normalized, 18.58× becomes 1.40×. Qwen3-8B's mean L2 norm is 17.4; Mistral's is 1.1 — that is the entire effect.
The probe accuracies above are computed independently and are unaffected. The corrected reading — near-identical geometry across labs — is a stronger claim than the one being withdrawn.
emotion-vector-bench/results/_comparison.json · probe_results.json ×5 · recomputed from raw_vectors.npz, 2026-08-17
characterization: residue is not a clean linear feature
six attempts, one wall
R1 Phase 1
rigor gates + 6 cells
all gates PASS; 3 of 6 cells beat facts
completed as scoped
R1 Phase 1.5
calibration gate
failed twice; judge VOID
scorer not calibratable
R1 forensic
3 gates
G1b on the bar, 5/10 seeds; reported, no artifact
held-out event never run
R2
3 conditions
2/3 → SUGGESTIVE; then retracted twice
grader defect, then fidelity confound
R3
5 gates
eval built, 1,265 items; 0/5 gates
hardware ceiling
emotion-vector-bench
replication
replicates on 5 models; headline corrected
completed
Read the fourth column. In four cases a pre-registered rule said stop and it was obeyed. In one case the instrument was characterized and the programme changed instruments. In two cases the work is genuinely unfinished, and both are named above rather than described as parked.