Anthropic showed that emotion concepts live as linear directions in a model's residual stream. They proved it on closed Sonnet 4.5. This reproduces the methodology on five open-weight models from three labs, in one command, on consumer hardware.
The finding: emotion geometry is close to model-invariant.
Across three labs and a five-fold parameter range, every geometric property lands in a narrow band. A 1.5B model comes within 4.6 points of an 8B on 20-way emotion classification. The differences between model families are smaller than the differences between layers inside one model.
| model | lab | probe acc | × chance | valence sep | effect d | cohesion | layer stab. |
|---|---|---|---|---|---|---|---|
| Llama-3.1-8B-Instruct | Meta | 91.5% ±0.9 | 18.3× | 0.716 | 1.36 | 0.252 | 0.987 |
| Mistral-7B-Instruct-v0.3 | Mistral | 91.3% ±0.8 | 18.3× | 0.693 | 1.30 | 0.247 | 0.987 |
| Qwen2.5-7B-Instruct | Alibaba | 90.5% ±1.0 | 18.1× | 0.694 | 1.22 | 0.215 | 0.961 |
| Qwen3-8B | Alibaba | 90.2% ±0.6 | 18.0× | 0.671 | 1.24 | 0.233 | 0.980 |
| Qwen2.5-1.5B-Instruct | Alibaba | 86.9% ±1.3 | 17.4× | 0.636 | 1.13 | 0.222 | 0.962 |
Probe accuracy is 20-way emotion classification from a single residual-stream direction, mean across four probed layers ± standard deviation; chance is 5%. Valence sep is PC1 separation between positive and negative emotion clusters on unit-normalized vectors. Effect d expresses that separation in PC1 standard deviations, so it is comparable across models. Cohesion is within-cluster minus cross-cluster cosine similarity. Permutation tests return p < 0.001 for every model at every layer.
Withdrawn 2026-08-17: the “18× spread in valence-axis strength.”
Until today this project's published headline claimed Qwen3-8B had an 18× cleaner valence axis than Mistral, and concluded there were “two valid geometric profiles” — a Qwen family organized by a dominant valence axis, and a Llama/Mistral family organized by tighter distributed clusters. Both claims are withdrawn. The repository and its findings document have been corrected in place.
The defect was in check_pca. It ran PCA on the raw mean-difference vectors and reported PC1 separation in those raw units. Activation magnitude varies by an order of magnitude across model families, so the number was measuring vector scale, not geometry.
The test that settles it — correlate the reported separation against each model's mean vector L2 norm:
model raw PC1 sep mean L2 norm llama-3.1-8b-instruct 2.709 1.865 mistral-7b-instruct-v0.3 1.571 1.110 qwen2.5-1.5b-instruct 7.300 6.115 qwen2.5-7b-instruct 12.056 9.209 qwen3-8b 29.189 17.438 corr(raw PC1 sep, mean L2 norm) = 0.9896 R² = 0.979
The ordering of the five models by raw separation is identical to their ordering by activation magnitude. The apparent “scale ladder” inside the Qwen family — 7.3 → 12.1 → 29.2 — tracks the norms 6.1 → 9.2 → 17.4 almost exactly. It was never a ladder of valence organization.
Unit-normalize the vectors before PCA, so the comparison is between directions rather than magnitudes, and the spread collapses:
| measure | as published | corrected |
|---|---|---|
| valence spread, raw units | 18.58× | — |
| valence spread, unit-normalized | — | 1.13× |
| models passing the criterion | 5 / 5 | 5 / 5 |
| “two geometric profiles” | claimed | withdrawn |
What the second withdrawal rests on. The two-profiles claim combined the valence artifact with a real cohesion difference — Llama and Mistral at 0.247–0.252 against the Qwen family at 0.215–0.233. Remove the artifact and what remains is a 1.18× gap whose confidence intervals overlap for four of the five models at n=50. That is worth a sentence, not a taxonomy.
What is unaffected. Every probe accuracy, every permutation test, cross-layer stability, and the arousal and implicit-emotion results are computed independently of PCA scale and stand unchanged. The fix is in code/validate.py; the old value is retained as valence_separation_pc1_raw with a comment stating it must not be compared across models.
The corrected claim is stronger than the withdrawn one. “One model family organizes emotion 18× more cleanly” is a curiosity about Qwen. “Emotion geometry is close to invariant across labs and scales” is a reason to trust that a result found on one open-weight model transfers to another — which is what makes the methodology viable on consumer hardware.
Six stages: extraction → vectors → validation → probes → arousal → implicit emotion. Twenty emotions grouped into valence clusters, a frozen 3,050-item corpus, four probed layers per model, permutation baselines at every stage. One command per model. All figures were produced locally on consumer hardware (Apple M-series, 24 GB memory).
The corpus is frozen precisely so a later run can be compared to an earlier one. The correction above was found by recomputing from the stored raw_vectors.npz files rather than from any summary — which is the only reason it was findable three months later.