Every transformer runs the same prompt through its own geometry. Two models can read the same sentence and produce different lists of numbers. Those lists locate the sentence in each model’s internal space, which has many coordinates and no direct English interpretation.
The question this project asks is narrow and concrete:
Can the effective rank of a target model’s representations predict how hard it is to reconstruct those representations on held-out prompts from another model, after controlling for width, size, model family, and a matched random-source floor?
The answer, in this frozen 16-model panel: the preregistered association survives. Higher target effective rank is associated with harder held-out reconstruction, the coefficient is positive under every control and in every leave-one-family-out fit, and real-source reconstruction beat a matched Gaussian-source floor on all 240 directed pairs.
This article explains what those words mean, how the panel was built, and exactly what the result does and does not support. If “representation,” “FUV,” “effective rank,” or “Gaussian-source floor” is unfamiliar, the plain-language glossary defines each term with a small example.
First, a room full of sentences
The representations used here have between 896 and 3,072 coordinates. The illustration below compresses that idea into two dimensions. Each dot is one prompt. Matching labels mean both models read the same prompt. The two arrangements are different, but they are not arbitrary: prompts that carry related information preserve enough shared structure for a simple linear map to convert one arrangement into the other.
The tensor for one model is the table of prompt vectors — one row per prompt, one column per coordinate. It is called a hidden-state tensor because it is extracted from the model’s final layer, after the model has read each prompt to its end. The panel pools each prompt’s token vectors with a masked mean: average only the positions the model actually read, ignore the padding. One prompt becomes one row.
A two-dimensional illustration
Select a prompt to see where each model places the same sentence.
“Translation” here means one thing only: a fitted linear map that reconstructs the target model’s held-out prompt vectors from the source model’s vectors. It is linear reconstruction, never knowledge transfer, semantic equivalence, or any claim about what either model “knows.” A directed pair (source, target) is a different task from (target, source): the map is learned in one direction, and the reverse direction needs its own fit.
Learn on some prompts, test on different ones
The translator learns from a training set of prompts. It is then frozen and tested on held-out prompts that could not influence the fit or its regularization. Held-out scoring is what distinguishes reconstruction from memorization: the score measures how well the fitted map predicts representations for prompts it never saw.
The score is calculated only on the held-out quarter.
The score is FUV, the fraction of unexplained variance: prediction error divided by the error from always predicting the average. Lower is better. Zero is perfect reconstruction, one means the map is no better than the average guess, above one is worse than that guess.
Move the FUV score
Drag the control. The only rule you need: lower is better.
Better than always guessing the average. Lower is better.
What would failure look like?
A low FUV on its own is hard to interpret, because any reconstruction pipeline enjoys some baseline success from sample size and dimensionality alone. To measure the structure associated with correct prompt alignment, the experiment breaks the alignment deliberately: keep both models’ real geometry, but pair each source prompt with the wrong target prompt.
This row-permutation baseline preserves the source tensor’s exact scale, spectrum, and geometry while destroying the prompt-to-prompt correspondence. The gap between the real score and the shuffled score is the evidence for prompt-aligned structure. In the panel this intuition is pushed one step further, with random-number source tensors replacing the source model entirely — the Gaussian-source floor described below.
Show me the FUV math
For a held-out target matrix $Y$ and reconstruction $\hat{Y}$:
$$ \operatorname{FUV}(\hat{Y},Y) = \frac{\sum_{i,j}(\hat{Y}_{ij}-Y_{ij})^2} {\sum_{i,j}(Y_{ij}-\bar{Y}_{\cdot j})^2}. $$
The clue: three models, two prompt sets
The study did not start at 16 models. An earlier experiment used three checkpoints — Qwen2.5-0.5B, TinyLlama-1.1B, and SmolLM2-1.7B — and two separately frozen sets of 64 prompts. All six directed maps beat the shuffled baseline on every one of 20 held-out splits in both fixtures, and the two pairings that involve TinyLlama were directionally asymmetric. The chart below is that historical result; every line runs from the real score to the shuffled score, farther left is better.
Median real FUV0.818
Median advantage over shuffle0.285
Splits won20 of 20
Every direction beat the shuffled baseline on all 20 splits in both fixtures. These numbers describe linear translatability between three frozen checkpoints. They do not compare the models’ intelligence or overall quality, and they cannot generalize: three checkpoints share prompt sets, widths, and lineage in ways that make a spurious common-cause explanation plausible.
This is where the earlier study stopped. It is the clue, not the endpoint. It suggested a hypothesis about the target: models whose variation is spread across more directions should be harder to reconstruct, because the fitted map must predict more independent coordinates. Testing that hypothesis across a family of models is what the frozen panel does.
RankMe: how spread out is the target’s variation?
Effective rank is a single number summarizing how much of a tensor’s variation is concentrated in a few directions versus spread across many. The panel uses RankMe, computed from the singular values of the model’s centered hidden-state tensor:
with singular values kept above times the largest one. RankMe is computed once per model on the full 4,000-row float32 tensor, with epsilon 1e-15, exactly as frozen in the design. A tensor whose energy sits in one direction has RankMe near 1; a tensor with energy spread evenly across many directions has RankMe near the number of directions.
The lab below shows the idea with eight synthetic directions. It is illustrative only: the eigenvalues are drawn, not measured. The total energy is identical at every slider position — the spectrum is only reshaped, concentrated or spread.
Effective rank: concentrated versus spread
Illustrative synthetic eigenvalues — not panel data
Effective rank ≈ 4.93 of 8 directions.
Slide left to concentrate the same total energy into a few directions, right to spread it evenly. The readout is the RankMe of the displayed spectrum. If all energy sat in one direction, effective rank would be 1.
Effective rank is a descriptor of the target tensor, not a count of concepts, not a quality score. It is the predictor this project asks about: does a target with more spread-out variation reconstruct worse, all else held fixed?
The frozen panel
The panel freezes every numerical choice before any real tensor is scored. That freeze is recorded in the repository design manifest and amendment, and it is what makes the result testable rather than post-hoc.
- 16 models, 12 families, all public instruction-tuned checkpoints between 0.49B and 3.83B parameters, selected before extraction. Every family contributes at least one model.
- 4 batteries × 1,000 prompts: general instruction, multiple-choice reasoning, grade-school math, and human conversation. One hidden-state tensor per model, 4,000 rows (4 × 1,000), extracted from the last layer with masked-mean pooling over each prompt’s tokens.
- 10 fixed split seeds. Each seed permutes the 1,000 rows of each battery; the first 200 rows are held out, the remaining 800 are the ordered training pool.
- Nested training sizes n = 64, 128, 256, 512, 800: for each n, training is exactly the first n rows of the same ordered pool. The held-out 200 never enter training at any n.
- Ridge with train-only CV: five-fold cross-validation inside the training rows selects the penalty; the untouched holdout is scored exactly once. Fold assignment is a deterministic hash of the canonical row id, so it is reproducible.
- 240 directed pairs: every ordered (source, target) pair among the 16 models, 16 × 15. Each pair’s outcome is the mean of its 40 nuisance cells — 4 batteries × 10 split seeds — at n = 800.
| Model | Family | Width | Parameters | RankMe |
|---|---|---|---|---|
| Llama-3.2-1B | llama | 2,048 | 1.24B | 1332.2 |
| Llama-3.2-3B | llama | 3,072 | 3.21B | 1845.0 |
| Qwen2.5-0.5B | qwen | 896 | 0.49B | 537.6 |
| Qwen2.5-1.5B | qwen | 1,536 | 1.54B | 884.1 |
| Qwen2.5-3B | qwen | 2,048 | 3.09B | 1202.8 |
| Phi-3.5-mini | phi | 3,072 | 3.82B | 1968.6 |
| Phi-3-mini | phi | 3,072 | 3.82B | 1938.0 |
| SmolLM2-1.7B | smollm | 2,048 | 1.71B | 1010.1 |
| Falcon3-3B | falcon | 3,072 | 3.23B | 1772.0 |
| OLMo2-1B | olmo | 2,048 | 1.48B | 1408.7 |
| Granite-3.1-2B | granite | 2,048 | 2.53B | 1367.0 |
| Gemma-2-2B | gemma | 2,304 | 2.61B | 1313.4 |
| TinyLlama-1.1B | tinyllama | 2,048 | 1.10B | 1331.7 |
| StableLM2-1.6B | stablelm | 2,048 | 1.64B | 1234.8 |
| OpenELM-3B | openelm | 3,072 | 3.04B | 1394.1 |
| BLOOMZ-3B | bloom | 2,560 | 3.00B | 1043.8 |
RankMe values shown to one decimal place; exact values live in the machine-readable summary.
The 240 directed pairs are not 240 independent experiments. Every pair shares one of 16 tensors as its source or its target, and all pairs are scored on the same batteries, the same split seeds, and the same regularization. The effective sample unit is the frozen panel of 16 models (12 families). Pair, battery, and seed counts are precision and stability detail; they are not independent replicates. This is the single most important constraint on how the results below may be read.
H3: real tensors against a matched random floor
Before the rank hypothesis, a sanity check: do the real tensors reconstruct each other any better than random numbers with the same shape and the same pipeline?
For every directed pair and every source width, battery, and split seed, the pipeline is rerun with a Gaussian random source tensor in place of the real source, drawn from a fixed inner seed and scored with the same CV-selected ridge. Each pair’s matched floor is the mean over the 20 source seeds of its Gaussian FUV at n = 800. The difference, observed minus Gaussian, is negative when real-source reconstruction was easier than the matched random floor.
The 16 × 16 matrix below shows that difference for all 240 directed pairs. Diagonal cells are unavailable — a model never reconstructs itself. Every off-diagonal value is negative. Focus, hover, or tap a cell to see the source, the target, the exact difference, and what it means; arrow keys move between cells when one is focused.
All 240 directed pairs: observed minus matched Gaussian-source FUV
Negative means real-source reconstruction beat the matched random floor
Source—
Target—
Difference—
Focus, hover, or tap a cell to inspect that directed pair.
Global statistic over all 240 pairs
One-sided add-one p = 1/21 ≈ 0.047619, Holm-adjusted p = 0.047619 (family size 1, alpha 0.05). Twenty matched source seeds; all 20 seed-level Gaussian panel means fall between 1.006592 and 1.006699, and none is at or below the observed mean.
Two features of this figure deserve emphasis. First, the p-value resolution is coarse by design: with 20 matched source seeds, the add-one p-value can only take values in steps of 1/21, and the observed value is the smallest possible nonzero step. A stronger test would need more matched seeds. Second, the statistic is a comparison of panel means. It says the real tensors reconstruct each other below a matched random floor; it says nothing about semantics, about which coordinates matter, or about why the real tensors are easier.
H1: does target effective rank predict difficulty?
H3 establishes the floor; the primary hypothesis asks whether the target’s effective rank predicts the difficulty of reconstructing it. The model is a ridge regression over the 240 aggregated pair rows:
Show the fitted formula
mean_fuv ~ z(log(target_rankme)) + z(log(source_rankme))
+ z(log(target_ambient_dimension)) + z(log(source_ambient_dimension))
+ z(log(target_parameter_count)) + z(log(source_parameter_count))
+ z(gaussian_source_floor_mean) + same_family
Each log predictor is standardized (z-scored). The reported coefficient is on the standardized log target RankMe.
The controls, in plain terms: the source’s own effective rank, the source and target ambient widths, the source and target parameter counts, the matched Gaussian floor for the pair, and an indicator for same-family pairs. The coefficient on the target’s log RankMe at n = 800 is
0.1811945678,
positive, with an add-one family-wild p-value of 0.0000199996 from 50,000 permutations (seed 314159) in which zero draws exceeded the observed statistic. Leave-one-family-out refits the model 12 times, dropping each family in turn; all 12 coefficients are positive (12/12), so the association is not carried by any single family. The frozen decision rule — positive coefficient, wild-null p below 0.05, at least 80% of leave-one-family-out coefficients positive — returns association_survives.
What the coefficient means: because the predictor is standardized, one standard deviation higher log target RankMe is associated with about +0.181 higher mean held-out FUV after the frozen controls. Higher FUV is harder reconstruction. So the association runs in the direction the clue suggested: targets whose variation is spread across more directions are reconstructed less well, holding width, size, family, source properties, and the random floor fixed.
The two views below show the same coefficient under two different lenses. The training-size view sweeps n = 64, 128, 256, 512, 800; n = 256/512/800 are the confirmatory sign checks, n = 64/128 are diagnostic only. The family view drops one family at a time. Both views are keyboard accessible: use the tabs to switch, then read the bars.
Stability of the target-RankMe coefficient
n = 640.2079096868diagnostic
n = 1280.2020017506diagnostic
n = 2560.1813243368confirmatory
n = 5120.2027746149confirmatory
n = 8000.1811945678confirmatory
drop bloom0.2331383378
drop falcon0.1801966086
drop gemma0.1836107492
drop granite0.1894152324
drop llama0.1847383917
drop olmo0.1864593511
drop openelm0.1853757020
drop phi0.1148879727
drop qwen0.1175163605
drop smollm0.1881392567
drop stablelm0.1855003207
drop tinyllama0.1883412510
Coefficients shown to ten decimal places; bar lengths are relative to the largest coefficient in each view. All 17 target-RankMe coefficients are positive.
H2 and H4: threshold rule and sign stability
Two preregistered companions sharpen the primary result without adding new p-values.
H2 is a threshold rule: at least 80% of estimable leave-one-family-out target-RankMe coefficients are positive. The observed fraction is 12/12, so the rule passes. It is a threshold, not a significance test.
H4 checks the coefficient’s sign across training sizes: positive at n = 64, 128, 256, 512, and 800, with n = 256/512/800 confirmatory sign checks and n = 64/128 diagnostic only. The sign is positive everywhere; H4 passes. No p-value is invented for either rule.
H5 and H6: mechanical checks on the pipeline
Two sentinel controls run on six exact directed pairs (both directions of llama32_1b ↔ llama32_3b, qwen25_05b ↔ phi35_mini, and smollm2_17b ↔ granite31_2b) at one frozen cell: battery general instruction, first split seed, n = 800. They are deterministic diagnostics with frozen tolerances. They are not significance tests, and no p-value is attached to them.
H5 is an invertibility check. The target’s spectrum is edited in a way that is exactly invertible — the edit is applied, then mapped back to raw coordinates — so the raw-space reconstruction should be unchanged. If the pipeline were corrupting predictions along the way, the raw predictions would move. The tolerance is a maximum absolute raw-prediction difference of 1e-9. All 6 pairs pass; the largest difference across all of them is 1.389821591146756e-11. The edited-space FUV changes (the edit is real), but the raw-space predictions do not.
H6 is a truncation floor. The target is projected onto its top-k training components, retaining at least 90% of training variance (k chosen as the smallest such count), and the mechanical FUV of that truncation is the oracle floor. The scored raw FUV of the ridge pipeline must not fall below that floor by more than 1e-10. All 6 pairs pass. This is a lower-bound check: the scored FUV sits comfortably above the floor in every pair, and the rule is that it may not dip below it — not that it lands within 1e-10 of it.
Both controls exist to catch pipeline failures, not to establish mechanism. H5 is a raw-space invariance negative control; it says nothing about what information the edit removed. H6 is a numerical sanity bound.
Reproducibility and audit
The analysis ran to completion exactly as frozen, and a separate read-only auditor re-derived what it could without rerunning the 27,206-cell computation.
| Quantity | Value |
|---|---|
| Real cells (16 sources × 4 batteries × 10 split seeds × 5 training sizes) | 3,200 |
| Gaussian cells (6 unique source widths × 20 source seeds × 4 batteries × 10 split seeds × 5 training sizes) | 24,000 |
| Sentinel cells (H5/H6) | 6 |
| Checkpoints total (real + Gaussian + sentinel + RankMe + design seal) | 27,208 |
| Ridge fits (2,208,000 real + 17,664,000 Gaussian + 552 sentinel) | 19,872,552 |
| Resume determinism | second run and resume produced an identical canonical payload (SHA-256 match) |
| Independent audit status | VERIFIED_COMPLETE (2026-08-08); zero failures |
| Row-hash comparisons recomputed by the auditor | 54,692 / 54,692 |
| Test suite at the frozen commit | 614 tests, warnings as errors, compilation passed |
| Scientific-core branch coverage | 99% (preregistered gate ≥ 95% met); whole-package branch coverage is 92% and is reported as-is — it does not reach the package-wide fail-under=95 threshold and is not a passed gate |
| Prime teardown | zero active pods and zero disks after the runs |
The 16 model tensors, the analysis checkpoints, and the audit report are external runtime artifacts and are not committed to the repository. What is committed is the frozen design, the code, the tests, and the machine-readable results. The intended reading order for public readers, all in the repository on main:
- results/panel_analysis_v1.summary.json — the compact publication summary.
- results/README.md — the plain-language result.
- docs/panel-analysis-postresult-audit-methodology.md — what the independent audit did.
- configs/panel-analysis-design.manifest.json — the sealed design manifest.
What this does not claim
The result is a panel-level association, and the boundaries are part of the result.
- No universal law. Sixteen models across twelve families are one frozen panel, not a sample from the space of all language models. Nothing here scales the association to other checkpoints, other layers, or other pooling rules.
- No causation. The target-spectrum controls adjust, they do not intervene. The coefficient says the association survives the controls; it does not say effective rank causes difficulty.
- No independent evidence from the 240 pairs. Directed pairs share endpoints, batteries, and seeds. The effective sample unit is the panel, and every inferential statement above respects that.
- No semantics and no knowledge transfer. Reconstruction is linear prediction in a hidden-state space. Lower FUV on held-out prompts says nothing about what either model knows, and “translation” is a geometric operation, not a meaning-preserving one.
- No downstream utility and no model-quality ranking. A target that is easier to reconstruct is not a better or worse model; the direction of the association is about geometry, and the earlier three-model asymmetry already showed that translatability does not track size or quality.
The panel answers the question it froze: after fixed controls, target effective rank predicts held-out reconstruction difficulty, and real tensors reconstruct each other below a matched random floor. That is the completed experiment, stated as it stands.