The question posed by the empirical programme
The observation of latent alignment and the need for a distinct ULR object are different claims. The empirical programme was designed to close the following chain:
1. The first six empirical atoms
| ID | Current observation | Permitted interpretation | Prohibited interpretation |
|---|---|---|---|
ULR-CONV-EXIST | Relational geometry exceeds a shuffled null for a limited set of concept/model pairs | Shared alignment exists for those pairs | Observer-independent common space |
ULR-CONV-INDEP | Independent vision-only × language-only pairs also exceed chance | The phenomenon is not solely a product of joint CLIP training | A common ULR for every modality |
ULR-CONV-DIM | Axes exceeding the held-out null appear in 20/25 CIFAR dimensions and 24/25 Caltech dimensions | The phenomenon is not confined to one axis | Exact intrinsic dimension 17, 20, or 24 |
ULR-CONV-SEM | Alignment has a graded association with semantic/language proxies | Association with the selected proxies | Identity with meaning |
ULR-CONV-ORIGIN | A typed trend of develops during training from a near-chance initialisation | The statistic strengthens during learning | An exact onset of formation |
ULR-EMP-GAUGE-ROBUST | The T1/T2 exact results and T3 trend are preserved under the specified permutation × positive-diagonal typing | The phenomenon is not merely a raw-coordinate artefact | Robustness under general GL, nonlinear, or full-model gauges |
These six rows are not six discoveries of ULR objects. They are active claims recording phenomena and artefact boundaries.
2. Quantities confirmed by the zero-assumption reopening
The reopening audit did not simply reuse earlier headlines. It separated sources, nulls, and model families again.
| Test | Result | Canonical interpretation |
|---|---|---|
| Raw cross-modal CKA | 0.769 | Relational statistic for the specified model pair |
| Shuffled baseline | 0.583 | Null under the same protocol |
| Standardised gap | Above chance for this pair | |
| Supervised × supervised, 14-model atlas | 0.803 | High alignment within an objective family |
| Self-supervised × self-supervised | 0.468 | Inconsistent with universal convergence |
| Untrained × untrained | 0.571 | Architecture and data priors form a non-negligible baseline |
| Variance explained by graded animacy proxies | Approximately 25% | Explanatory power of the selected proxy family |
| Unexplained variance | Approximately 75% | Residual unexplained by 18 proxies, not the size of ULR |
Earlier numbers lacking sufficient provenance—including ≥17, shared/private 52%, 48% after
principal-component removal, and an exact formation onset—carry no ontological weight in the current canon.
3. How gauge-typed replication changed the conclusion
A raw representational metric can change under a function-preserving coordinate transformation. The T1–T3 replications compared raw-coordinate measurements with their quotient-normalised, gauge-typed counterparts.
- In T1, raw effective dimension changed from 20 to 16, while the typed value remained bit-identical.
- In T2, the raw modality-specific score fell from 0.251 to 0.119, while the typed score remained 0.237.
- T3 initially failed to replicate. After the discrepancy was traced to the extraction protocol, the original trend was recovered under the last-token, final-layer, fixed-prompt condition, with CKA correlation 0.997 and retrieval Pearson correlation 0.986.
The conclusion was frozen in one sentence:
The phenomenon was not a gauge artefact, but the raw measurement was exposed to gauge artefacts.
This does not mean that the entire full-model gauge orbit was tested. The result is limited to the specified permutation and positive-diagonal typing.
4. B0–B4 and combined B*
The value of a new quantity is not tested against one weak baseline, but against the strongest available explanation under the same information and port contract.
| Baseline | Observation scope | Misreading it is designed to remove |
|---|---|---|
| B0 output | Passive output only | Calling an output difference a new internal object |
| B1 geometry | Passive representational geometry | Redundant information in a relational metric |
| B2 predictive | Passive future prediction | Absorption into ordinary predictive state |
| B3 causal testing | Raw controlled port response | Absorption of the response quotient into standard testing |
| B4 routing | Architecture-declared state and write routing | Mistaking a design label for emergent organisation |
| B* | B0–B4 combined | Vacuous promotion after separately beating several weak nulls |
To open M8, had to add incremental value for held-out behaviour or transfer after B*.
5. U0-v1 — Moving the exact toy into a small neural system
This experiment connected the same trained core to two different persistence mechanisms.
- A: preserve the write activation in recurrent hidden state.
- B: preserve it in a one-item attention cache and reconstruct the activation at query time.
- 32 seeds, with 24 for meta-training and 8 for meta-testing.
- Pairwise-exact activations and probabilities during passive write, tick, and read operations.
- Hidden/cache reset, clone, and overwrite operations act on the actual state container.
The maximum difference in passive probabilities and activations was 0, and CKA was 1. In 32/32 seeds, the recurrent-hidden-minus-attention-cache response contrast was negative under reset and positive under clone. Under signed-permutation co-transformation, passive output and physical response were also preserved to numerical precision.
Held-out prediction
| Predictor | Test MSE | Test |
|---|---|---|
| B0 output | 0.9652701411 | 0.0 |
| B1 passive geometry | 0.9652701411 | 0.0 |
| B2 passive predictive | 0.9652701411 | Approximately 0 |
| B3 raw port response | 0.0000105580 | 0.999989 |
| B4 declared routing | 0.0000014978 | 0.999998 |
| Q typed response | 0.0000007289 | 0.999999 |
substantially outperformed B2, but its advantage over B3 and B4 fell below the preregistered 0.01
nonabsorption threshold. The operational bridge is therefore positive, while the ontology verdict is
Z1_WEAKENED.
6. M1 — An emergent version of the Z1 physical response
Because U0-v1 used engineered routing, M1 tested whether stable response classes would emerge naturally from seed-only training of a homogeneous tanh RNN.
- All 64 seeds passed passive eligibility.
- Stable classes failed to form: the majority-to-outlier cluster counts were 47:1 for meta-train and 16:0 for meta-test.
- The raw response also failed to improve on an ordinary-dynamics baseline.
- Local verdict:
COLLAPSED@tanh-delay-v2.
This result does not refute every physical-response hypothesis. It closes the current implementation of emergent v2 while leaving Z1 as a weakened application programme.
7. M2 — Z5 external-relation cross-fit
M2 fixed an external semantic relation and tested whether a relation law fitted on one dataset transported to another dataset and model family.
- The fixed WordNet 3.1 relation predicted held-out Caltech and CIFAR relations.
- The result held across architecture and objective holdouts.
- The geometry-relation gate was positive.
- The retrieval-behaviour bridge deteriorated and was negative.
Z5 is therefore a weakened external baseline. The finding that an external taxonomy predicts some representational geometry survives, but it is not promoted to an internal carrier or behavioural mechanism.
8. M3 — A uniform reduction audit of seven carriers
| Candidate | Reduction target | Neural-specific survivor |
|---|---|---|
| Parameter quotient | Symmetry/invariant quotient | 0 |
| Activation ensemble | Empirical law/kernel/moment | 0 |
| Relation transform | RSA/CKA/external relation | 0 |
| Predictive state | PSR/minimal realisation | 0 |
| Physical response | Causal testing | 0 |
| Dynamical carrier | Ordinary state-space/RDS | 0 |
| External relation | Not an internal object | 0 |
The final count is 0/7. This verdict closes the current registry; it does not prohibit every possible future carrier candidate.
9. M4 — A conditional definition of formation
Formation was defined not as the onset of a metric curve, but as an event in which the structural type of an object or family changes under identity and transport.
The required elements are:
- object identity before and after formation;
- cross-time transport;
- a distinction between motion within a quotient and change of structural stratum;
- a separate behavioural bridge, not built into the definition;
- a distinction between deformation and boundary crossing.
Because M3 has zero survivors, the current target across neural models is NO_ADMISSIBLE_TARGET. The definition
is complete, but no universal instance has been established.
10. M5 — Retraction of the UAR composite
M5 re-examined the earlier impression that UAR-I had been completed. It found that injectivity of the full-channel incidence map and unramifiedness of the joint assembly map are propositions about different sources and maps.
Accordingly, ULR-ASSEMBLY-UAR-MONO was downgraded to RETRACTED-MALFORMED-COMPOSITE. Two subordinate
results were preserved:
ULR-UAR-CH-TYPED-EXCLUSION;ULR-UAR-ASM-UNRAMIFIED.
The localisation identity and CERT-B/C were moved to an optional strengthening layer that does not block M8.
11. M6 — Closing the baseline contract
M6 fixed B0–B4 and B* so that every subsequent candidate would be judged by the same standard. Candidate and baseline must share the same split, information access, intervention vocabulary, precision, and resource ledger. If a baseline can execute the candidate wrapper, excluding that execution requires a separate provenance-blind doctrine.
12. M7 — Closing the hypotheses and governance
A contract was completed that preregisters the following for each of Z0–Z7:
- object and identity;
- strongest rival;
- rival-specific falsifier;
- independent prediction;
- data, split, hash, and amendment rule;
- finite termination;
- raw receipt and validator;
- separation of author closure from independent audit.
The 14-field schemas of all eight hypotheses passed format validation. Passing a format check does not establish truth or novelty.
13. M8 — The final ontology verdict
M8 did not take a simple vote over the preceding results. It asked whether any residual remains after every typed baseline.
| Candidate basis | Observed result | M8 interpretation |
|---|---|---|
| Physical response | Richer than the passive quotient | Absorbed by B3/B4 testing and routing |
| Emergent persistence | No stable class | Z1 v2 collapsed |
| External relation | Positive geometry cross-fit | Negative behaviour; not an internal carrier |
| Carrier registry | Rich toolkit | Survivor count 0/7 |
| Formation | Object-relative definition available | No universal target |
| UAR | Two subordinate results survive | Composite monomorphism retracted |
Therefore,
was adopted. This NO is a conclusion within the finite registry and the declared baseline scope.
14. Canon 24's post-verdict theorem
Even after M8, the observer-relative role problem could still be characterised exactly. For role-conditioned transcript laws , the equal-prior Bayes risk is
Adaptive distinguishability is the supremum of total variation over admissible policies. Separately, data processing contracts total variation when an internal transcript is a garbling of an external transcript. Under a fixed infinite policy, finite-prefix error tends to zero if and only if the full path laws are mutually singular. These distinct results did not reopen M8: they establish limited-observer identifiability boundaries, not a new ontology.
15. Lessons from reproduction and audit
- Saved-output computations matched 13/13 stdout hashes in the current environment.
- The U0-v1 validator passed 12/12 checks covering preregistration, raw shape, split, gate, verdict, and deterministic replay.
- The finite verifier for the observer theorem passed 8/8 checks, including distribution pairs, garbling, and adaptive policies.
- Code reproduction does not replace the full input-provenance and external-access contract.
- Independent audit overturned an author-side self-audit finding of “not a result defect” and identified a conclusion-level defect.
The current empirical discipline therefore treats author-side closure, independent audit, and LEDGER promotion as distinct stages.
Final empirical conclusion
Strong alignment phenomena and exact boundary theorems exist. Yet on the current data, the measured gain over
B3/B4 was nonzero but below the preregistered nonabsorption threshold; no neural-specific ULR candidate
therefore retained preregistered meaningful incremental value beyond B*. The empirical programme did not fail: its
present achievement is that it maintained null tests demanding enough to overturn its preferred ontology and reach
a NO verdict.