Supplementary material · Additional analyses

Supplementary material

Robustness checks, assignment figures, and null tests.

01 / Chapter one

Assignment and robustness checks

Question

How sensitive are the assignments to the threshold, number of concepts, and composition of the concept set?

Result

We compare these choices through coverage, the resulting assignments, and the case-study responses.

Conclusion

These checks show where the comparison is stable and where the interpretation remains conditional on the assignment procedure.

Target responses across five cases for 5, 8, and 10 concepts. Special relativity at 5 concepts is hatched as undefined; the other SR bars are ranked first among 3 and 6 measurable concepts. Higgs bars show a known assignment artifact.

01 / Granularity

The response depends on the available concept set.

We repeat the analysis with 5, 8, and 10 expert-defined concepts. Bar heights show the target's response, max⁡(∣DI∣,∣DP∣)\max(|D_I|, |D_P|), at the published pivot (the year separating the before-and-after comparison), holding fkneef_{\mathrm{knee}} fixed. Labels give its rank among concepts with a measurable response, not among all concepts included: 1/3 means first among three measurable responses. Hatching marks an undefined response, not zero.

Interpretation. Special relativity ranks first among 3 measurable concepts at N = 8 and 6 at N = 10. At N = 5, removing the target leaves too few active concepts for the geometry, so the response and rank are undefined. The nominal Higgs response remains an assignment artifact, not historical evidence.

Methods and additional context

Each set contains the target and the first N − 1 context concepts in alphabetical order, so composition and count change together. Paper assignments are recomputed for each set. A concept can therefore receive papers at one N but none at another. Radiation/quantum enters the set at N = 8; spectroscopy and thermodynamics enter at N = 10.

A measurable response requires assigned papers and enough time windows in which both the original and the concept-removed geometry can be calculated. The implemented geometry needs at least two active concepts in a window, and the standardized before-and-after comparison needs at least two valid windows on each side of the pivot. No assigned papers and too few valid comparison windows are different reasons for an undefined response. Neither is treated as a measured zero or included in the ranking.

Special relativity, N = 5: gravitation has no assigned papers. Special relativity and electron theory have assigned papers, but removing either leaves too few active concepts to calculate the post-pivot geometry. The matched comparison retains seven pre-pivot and no post-pivot windows for special relativity, and three pre-pivot and no post-pivot windows for electron theory. Only aether optics and electrodynamics have defined responses; the target itself cannot be ranked.

Special relativity, N = 8: gravitation, instrumentation/measurement, and mechanics/time measurement have no assigned papers. Aether optics and electrodynamics each have one assigned paper, but their concept-removal comparisons retain only one valid pre-pivot window each (and 17 post-pivot windows), which is insufficient. The three ranked concepts are special relativity, electron theory, and radiation/quantum. The label 1/3 therefore means first among these three, not first among all eight configured concepts.

Special relativity, N = 10: gravitation, instrumentation/measurement, mechanics/time measurement, and radiation/quantum have no assigned papers under this configuration. The six ranked concepts are special relativity, aether optics, electrodynamics, electron theory, spectroscopy, and thermodynamics/kinetic theory. The label 1/6 means first among these six. Aether optics and electrodynamics now have enough valid comparison windows, although each still has only one assigned paper: measurability depends on the surrounding geometry and temporal coverage, not paper count alone.

Among ten sampled margin scales, special relativity never ranks first at N = 5 and ranks first at N = 8 only at f = 0.25. These are discrete samples, not a test of every margin scale or every possible concept subset. Gödel and deep learning rank first at N = 8 and 10; attention remains at ranks 3–4.

Assignment rule used in the paper: Jaccard overlap for single paraphrases, 1-versus-4 and 2-versus-3 splits, and leave-one-out anchors. Labels show reduced concept coverage; special relativity has only one concept for single-paraphrase comparisons versus five for 2-versus-3 splits, so those bars are not a matched comparison.

02 / Semantic stability

Does changing the description change which papers are selected?

If we describe the same scientific concept using different words, does the method select the same papers? We test five descriptions of each concept, both individually and in combinations. The vertical axis measures overlap between the selected paper lists: 0 means no shared papers and 1 means identical lists. Each case-study group summarizes multiple concepts, not just the target named on the horizontal axis. Fractions above the bars show how many of the ten concepts contribute usable comparisons; unlike the concept-count plot, these labels are coverage counts, not ranks.

Interpretation. For Higgs, Gödel, deep learning, and attention, combining two or three descriptions gives more consistent paper lists than using individual descriptions, when comparing the same measurable concepts. For special relativity, individual-description comparisons cover only 1 of 10 concepts, versus 5 when comparing combinations of two and three descriptions: those bars are not a matched comparison. Agreement tests sensitivity to wording, not whether the selected papers are historically correct.

Methods and additional context

Jaccard overlap is the number of papers shared by two lists divided by the number of distinct papers appearing in either list. For example, lists containing papers A, B, C and B, C, D share two of four distinct papers, giving an overlap of 0.5. This measures agreement between selections, not the fraction of scientifically correct assignments.

A paraphrase is one of the five descriptions of a concept. We combine descriptions by averaging their embedding vectors, then use the resulting vector to select papers. Orange bars compare one description with another. Blue bars compare one description with the other four combined. Black bars compare two descriptions combined with the remaining three combined; the two groups share no descriptions.

The hollow green bars compare four descriptions combined with all five. This asks how much the selection changes when one description is omitted. Because the two sides share four descriptions, high overlap here is not independent evidence that different descriptions produce the same assignments.

For each concept, we average the overlap over the usable comparisons. Bars show the median of those concept-level averages; whiskers show the interquartile range across concepts. Comparisons in which both lists are empty are excluded because they provide no selected papers to compare. If only one list is empty, the overlap is zero. A concept contributes to a bar only if at least one comparison remains.

Assignment rule used in the paper: a nearest-anchor-group 60th-percentile similarity threshold and a confusion-aware margin τ(c)\tau(c), with fkneef_{\mathrm{knee}} fixed per case. Only the anchor under study varies; the other nine retain their full five-paraphrase anchors. Thresholds and margins are recomputed for each variant.

The numbers in parentheses in the legend count comparisons per concept: up to 10 single-paraphrase pairs, 5 unordered 1-versus-4 splits, and 10 unordered 2-versus-3 splits. The fractions above bars instead count contributing concepts. For special relativity, these are 1/10 for orange, 6/10 for blue, 5/10 for black, and 6/10 for hollow green. The sole contributing concept for the orange bar is electron theory, not special relativity itself. Differences between these bars therefore combine differences in description choice and concept coverage; they are not a matched comparison across the same concepts.

The corrected comparison excludes empty-versus-empty pairs and counts each unordered split once, regardless of its original orientation. The earlier assignment-rule comparison remains in the downloadable source data; only the assignment rule used in the paper is shown here.

02 / Chapter two

Concept anchors and assignment geometry

Question

How do five written paraphrases become one anchor, and when is a paper assigned to it?

Result

Each anchor is the ℓ2\ell_2-normalized mean of five paraphrase embeddings. A paper is assigned to its closest anchor only if it clears both that concept's 0.60-quantile similarity threshold and its confusion-aware margin threshold.

Conclusion

The anchor combines multiple descriptions rather than using one phrase alone. The margin requirement rejects papers that are similarly close to competing concepts; averaging does not guarantee higher assignment overlap.

More details: You can find the complete 50-concept reference, including descriptions, aliases, and paraphrases, here.

PC1 and PC2. Panel (a) uses principal component analysis (PCA), not UMAP or t-SNE. PC1 and PC2 are the first and second principal components of the 50 paraphrase embeddings in each case. The ensemble anchors are projected into the same two-dimensional display without refitting the axes. PCA is used only for visualization; assignments are evaluated in the full embedding space.

Gray points in panel (b). This panel plots similarity to each paper's best-matching anchor against its margin over the runner-up; it is not a low-dimensional projection. Orange points are assigned to the target. Gray points include papers assigned to other concepts and papers labeled “no match.” The lines mark the target's thresholds; the shaded bands show the ranges for other concepts. Crossing both target lines does not imply target membership: the target must also be the paper's closest anchor, and other concepts have their own thresholds.

Special relativity: anchor ensemble and assignment

Two-panel special-relativity assignment diagram showing five paraphrase embeddings combined into an anchor and documents evaluated against similarity and margin thresholds.

Five paraphrases form one concept anchor; papers are retained only when they clear both the similarity threshold and the confusion-aware margin.

Higgs mechanism: anchor ensemble and assignment

Two-panel Higgs-mechanism assignment diagram showing the paraphrase ensemble and document positions relative to the two assignment criteria.

The same anchor-ensemble and two-threshold rule is applied without changing the assignment geometry between cases.

Gödel incompleteness: anchor ensemble and assignment

Two-panel Gödel-incompleteness assignment diagram showing five paraphrases averaged into an anchor and document acceptance by similarity and margin.

The anchor combines five hand-written descriptions, while the margin rejects documents that are similarly close to competing concepts.

Deep learning: anchor ensemble and assignment

Two-panel deep-learning assignment diagram showing paraphrase embeddings, their mean anchor, and documents separated by similarity and margin thresholds.

Assignment depends on both absolute similarity and separation from the next-closest anchor.

Attention mechanism: anchor ensemble and assignment

Two-panel attention-mechanism assignment diagram showing the five-paraphrase anchor ensemble and document acceptance under the similarity and margin criteria.

The common assignment rule makes the five case studies directly comparable while retaining case-specific thresholds.

03 / Chapter three

Two-dimensional null distributions

Question

Is each joint response larger than expected after controlling for ablated-set size and paper identity?

Result

Random removal controls set size; scrambled assignment controls which papers carry a label. The paper's co-dominance test counts null draws whose absolute deviations from the null mean are at least as large as the observation's on both DID_I and DPD_P. It is two-sided on each axis and does not require a covariance-matrix estimate.

Conclusion

The two nulls address different explanations for the observed response and are therefore reported separately.

Interpretation note. The Higgs response remains an assignment artifact even though it passes these null tests. Entries marked < 2e-4 indicate the finite-permutation reporting floor, not a zero probability.

Scroll horizontally to see all columns →

Joint co-dominance p-values under the two null designs.
CaseRandom-removal pcop_{\mathrm{co}}Scrambled-assignment pcop_{\mathrm{co}}
Special relativity< 2e-4< 2e-4
Higgs mechanism< 2e-4< 2e-4
Gödel incompleteness< 2e-40.0004
Deep learning0.0232< 2e-4
Attention mechanism0.00060.0156

Random-removal nulls

Five panels of random-removal null draws in the D_I and D_P plane, with co-dominance regions and observed case-study points shown as stars.

Random removal asks whether the observed response is larger than removing the same number of papers at random.

Scrambled-assignment nulls

Five panels of scrambled-assignment null draws in the D_I and D_P plane, with co-dominance regions and observed case-study points shown as stars.

Scrambled assignment asks whether the response depends on which papers were assigned to a concept.

Special-relativity random-removal reference

Special-relativity random-removal null cloud in the D_I and D_P plane with the observed point and joint co-dominance region.

This is the manuscript's focused random-removal panel retained for direct reference.

Special-relativity scrambled-assignment reference

Special-relativity scrambled-assignment null cloud in the D_I and D_P plane with the observed point and joint co-dominance region.

This is the manuscript's focused scrambled-assignment panel retained for direct reference.

04 / Chapter four

Look-elsewhere effect and ranking robustness

Question

Does the result survive scanning the full concept-year grid, and does the target ranking depend on the chosen summary statistic?

Result

Special relativity and Gödel rank first under all three summaries; the nominal first-place Higgs result remains an assignment artifact. Deep learning ranks first under the maximum norm and directional p2Dp_{\mathrm{2D}}, but second under ∥D∥2\|D\|_2 in a near-tie with graphical models. Attention remains non-leading. The supplementary p2Dp_{\mathrm{2D}} ranking uses the observation's direction on each axis, unlike the paper's primary two-sided co-dominance test.

Conclusion

Ranking is broadly similar across these summaries, but is not identical. Special relativity survives the grid-wide correction; the nominal Higgs result inherits the assignment artifact and carries no independent evidential weight.

Interpretation note. For Gödel, deep learning, and attention, the grid-wide maximum occurs away from both the target concept and target pivot. Their pglobalp_{\mathrm{global}} values test the most extreme cell anywhere in the grid, not the target cell.

Scroll horizontally to see all columns →

Grid-wide maximum response and look-elsewhere correction across each 10-concept by 11-year grid.
CaseGlobal max ∥D∥2\|D\|_2pglobalp_{\mathrm{global}}zMaximizing cellOn target?
Special relativity7.6200.00087.913special_relativity, 1902Yes
Higgs mechanism14.9280.02022.675higgs_mechanism, 1959Yes
Gödel incompleteness2.8480.15080.869type_theory, 1940No
Deep learning20.6270.08821.218feature_engineering, 2008No
Attention mechanism6.0230.9646-1.154reinforcement_learning, 2008No

Scroll horizontally to see all columns →

Target ranking under the manuscript's maximum norm, joint Euclidean magnitude, and sign-aware p-value.
Caset∗t^*max⁡(∣DI∣,∣DP∣)\max(|D_I|, |D_P|)Rank∥D∥2\|D\|_2Rankp2Dp_{\mathrm{2D}}Rank
Special relativity19027.61217.6201< 2e-41
Higgs mechanism195914.569114.9281< 2e-41
Gödel incompleteness19372.49612.54410.00041
Deep learning20113.12713.4882< 2e-41
Attention mechanism20121.64041.87040.01583

Deep learning is rank 2 under ∥D∥2\|D\|_2 at 3.488, a near-tie with rank 1 graphical_models at 3.544. Attention mechanism improves from rank 4 under max⁡(∣DI∣,∣DP∣)\max(|D_I|, |D_P|) to rank 3 under p2Dp_{\mathrm{2D}} (4 → 3).

Grid-wide look-elsewhere correction

Five look-elsewhere null distributions of the maximum joint magnitude over each concept-year grid, with the observed maximum and its concept-year annotation.

The corrected comparison uses the maximum over the entire 10-concept by 11-year grid in both observed and permuted data.

Special-relativity look-elsewhere reference

Special-relativity look-elsewhere null distribution for the grid-wide maximum joint magnitude with the observed target maximum marked.

The manuscript's focused panel shows the special-relativity grid-wide correction.

05 / Chapter five

Target concept versus surrounding concepts

Question

Does a target exceed the typical surrounding concept, and does it also exceed the strongest competitor?

Result

The mean of nine contexts and the maximum single context are deliberately shown in separate panels on shared log-log axes.

Conclusion

The first panel compares the target with typical context. The second applies the stricter comparison with the strongest competitor.

Interpretation note. The Higgs point is retained to show the nominal measurement. Its large response is caused by pre-pivot sparsity and a misassigned paper, not evidence for a stronger conceptual reorganization.

Target response versus typical and strongest context

Two log-log scatter panels comparing each target response with the mean of nine context concepts and with the strongest single context concept; both include a y equals x line.

The left panel asks whether a target exceeds typical context; the right asks whether it exceeds its strongest competitor.