Skip to main content

ΩULR · Canonical· Canon 24

The ULR Empirical Programme — From Observation to M1–M8

The ULR empirical programme did not promote positive alignment directly into an ontology. It tested the phenomenon against gauges, nulls, untrained models, behavioural bridges, and B0–B4/B*. At the end of M1–M8, the carrier-survivor count is 0/7 and the ontology verdict is NO.

1,834 words7 min read

The question posed by the empirical programme

The observation of latent alignment and the need for a distinct ULR object are different claims. The empirical programme was designed to close the following chain:

phenomenonnull robustnessgauge typingstrong baseline comparisonheld-out behaviournonreduction verdict.\text{phenomenon} \to \text{null robustness} \to \text{gauge typing} \to \text{strong baseline comparison} \to \text{held-out behaviour} \to \text{nonreduction verdict}.

1. The first six empirical atoms

IDCurrent observationPermitted interpretationProhibited interpretation
ULR-CONV-EXISTRelational geometry exceeds a shuffled null for a limited set of concept/model pairsShared alignment exists for those pairsObserver-independent common space
ULR-CONV-INDEPIndependent vision-only × language-only pairs also exceed chanceThe phenomenon is not solely a product of joint CLIP trainingA common ULR for every modality
ULR-CONV-DIMAxes exceeding the held-out null appear in 20/25 CIFAR dimensions and 24/25 Caltech dimensionsThe phenomenon is not confined to one axisExact intrinsic dimension 17, 20, or 24
ULR-CONV-SEMAlignment has a graded association with semantic/language proxiesAssociation with the selected proxiesIdentity with meaning
ULR-CONV-ORIGINA typed trend of ρ=0.881\rho=0.881 develops during training from a near-chance initialisationThe statistic strengthens during learningAn exact onset of formation
ULR-EMP-GAUGE-ROBUSTThe T1/T2 exact results and T3 trend are preserved under the specified permutation × positive-diagonal typingThe phenomenon is not merely a raw-coordinate artefactRobustness under general GL, nonlinear, or full-model gauges

These six rows are not six discoveries of ULR objects. They are active claims recording phenomena and artefact boundaries.

2. Quantities confirmed by the zero-assumption reopening

The reopening audit did not simply reuse earlier headlines. It separated sources, nulls, and model families again.

TestResultCanonical interpretation
Raw cross-modal CKA0.769Relational statistic for the specified model pair
Shuffled baseline0.583Null under the same protocol
Standardised gapz=35.9z=35.9Above chance for this pair
Supervised × supervised, 14-model atlas0.803High alignment within an objective family
Self-supervised × self-supervised0.468Inconsistent with universal convergence
Untrained × untrained0.571Architecture and data priors form a non-negligible baseline
Variance explained by graded animacy proxiesApproximately 25%Explanatory power of the selected proxy family
Unexplained varianceApproximately 75%Residual unexplained by 18 proxies, not the size of ULR

Earlier numbers lacking sufficient provenance—including ≥17, shared/private 52%, 48% after principal-component removal, and an exact formation onset—carry no ontological weight in the current canon.

3. How gauge-typed replication changed the conclusion

A raw representational metric can change under a function-preserving coordinate transformation. The T1–T3 replications compared raw-coordinate measurements with their quotient-normalised, gauge-typed counterparts.

  • In T1, raw effective dimension changed from 20 to 16, while the typed value remained bit-identical.
  • In T2, the raw modality-specific score fell from 0.251 to 0.119, while the typed score remained 0.237.
  • T3 initially failed to replicate. After the discrepancy was traced to the extraction protocol, the original trend was recovered under the last-token, final-layer, fixed-prompt condition, with CKA correlation 0.997 and retrieval Pearson correlation 0.986.

The conclusion was frozen in one sentence:

The phenomenon was not a gauge artefact, but the raw measurement was exposed to gauge artefacts.

This does not mean that the entire full-model gauge orbit was tested. The result is limited to the specified permutation and positive-diagonal typing.

4. B0–B4 and combined B*

The value of a new quantity QQ is not tested against one weak baseline, but against the strongest available explanation under the same information and port contract.

BaselineObservation scopeMisreading it is designed to remove
B0 outputPassive output onlyCalling an output difference a new internal object
B1 geometryPassive representational geometryRedundant information in a relational metric
B2 predictivePassive future predictionAbsorption into ordinary predictive state
B3 causal testingRaw controlled port responseAbsorption of the response quotient into standard testing
B4 routingArchitecture-declared state and write routingMistaking a design label for emergent organisation
B*B0–B4 combinedVacuous promotion after separately beating several weak nulls

To open M8, QQ had to add incremental value for held-out behaviour or transfer after B*.

5. U0-v1 — Moving the exact toy into a small neural system

This experiment connected the same trained core to two different persistence mechanisms.

  • A: preserve the write activation in recurrent hidden state.
  • B: preserve it in a one-item attention cache and reconstruct the activation at query time.
  • 32 seeds, with 24 for meta-training and 8 for meta-testing.
  • Pairwise-exact activations and probabilities during passive write, tick, and read operations.
  • Hidden/cache reset, clone, and overwrite operations act on the actual state container.

The maximum difference in passive probabilities and activations was 0, and CKA was 1. In 32/32 seeds, the recurrent-hidden-minus-attention-cache response contrast was negative under reset and positive under clone. Under signed-permutation co-transformation, passive output and physical response were also preserved to numerical precision.

Held-out prediction

PredictorTest MSETest R2R^2
B0 output0.96527014110.0
B1 passive geometry0.96527014110.0
B2 passive predictive0.9652701411Approximately 0
B3 raw port response0.00001055800.999989
B4 declared routing0.00000149780.999998
Q typed response0.00000072890.999999

QQ substantially outperformed B2, but its advantage over B3 and B4 fell below the preregistered 0.01 nonabsorption threshold. The operational bridge is therefore positive, while the ontology verdict is Z1_WEAKENED.

6. M1 — An emergent version of the Z1 physical response

Because U0-v1 used engineered routing, M1 tested whether stable response classes would emerge naturally from seed-only training of a homogeneous tanh RNN.

  • All 64 seeds passed passive eligibility.
  • Stable classes failed to form: the majority-to-outlier cluster counts were 47:1 for meta-train and 16:0 for meta-test.
  • The raw response also failed to improve on an ordinary-dynamics baseline.
  • Local verdict: COLLAPSED@tanh-delay-v2.

This result does not refute every physical-response hypothesis. It closes the current implementation of emergent v2 while leaving Z1 as a weakened application programme.

7. M2 — Z5 external-relation cross-fit

M2 fixed an external semantic relation and tested whether a relation law fitted on one dataset transported to another dataset and model family.

  • The fixed WordNet 3.1 relation predicted held-out Caltech and CIFAR relations.
  • The result held across architecture and objective holdouts.
  • The geometry-relation gate was positive.
  • The retrieval-behaviour bridge deteriorated and was negative.

Z5 is therefore a weakened external baseline. The finding that an external taxonomy predicts some representational geometry survives, but it is not promoted to an internal carrier or behavioural mechanism.

8. M3 — A uniform reduction audit of seven carriers

CandidateReduction targetNeural-specific survivor
Parameter quotientSymmetry/invariant quotient0
Activation ensembleEmpirical law/kernel/moment0
Relation transformRSA/CKA/external relation0
Predictive statePSR/minimal realisation0
Physical responseCausal testing0
Dynamical carrierOrdinary state-space/RDS0
External relationNot an internal object0

The final count is 0/7. This verdict closes the current registry; it does not prohibit every possible future carrier candidate.

9. M4 — A conditional definition of formation

Formation was defined not as the onset of a metric curve, but as an event in which the structural type of an object or family changes under identity and transport.

The required elements are:

  1. object identity before and after formation;
  2. cross-time transport;
  3. a distinction between motion within a quotient and change of structural stratum;
  4. a separate behavioural bridge, not built into the definition;
  5. a distinction between deformation and boundary crossing.

Because M3 has zero survivors, the current target across neural models is NO_ADMISSIBLE_TARGET. The definition is complete, but no universal instance has been established.

10. M5 — Retraction of the UAR composite

M5 re-examined the earlier impression that UAR-I had been completed. It found that injectivity of the full-channel incidence map and unramifiedness of the joint assembly map aasm ⁣:Θadm/GΣJa_{\mathrm{asm}}\colon \Theta^{\mathrm{adm}}/G_\Sigma\to\mathcal J are propositions about different sources and maps.

Accordingly, ULR-ASSEMBLY-UAR-MONO was downgraded to RETRACTED-MALFORMED-COMPOSITE. Two subordinate results were preserved:

  • ULR-UAR-CH-TYPED-EXCLUSION;
  • ULR-UAR-ASM-UNRAMIFIED.

The localisation identity and CERT-B/C were moved to an optional strengthening layer that does not block M8.

11. M6 — Closing the baseline contract

M6 fixed B0–B4 and B* so that every subsequent candidate would be judged by the same standard. Candidate and baseline must share the same split, information access, intervention vocabulary, precision, and resource ledger. If a baseline can execute the candidate wrapper, excluding that execution requires a separate provenance-blind doctrine.

12. M7 — Closing the hypotheses and governance

A contract was completed that preregisters the following for each of Z0–Z7:

  • object and identity;
  • strongest rival;
  • rival-specific falsifier;
  • independent prediction;
  • data, split, hash, and amendment rule;
  • finite termination;
  • raw receipt and validator;
  • separation of author closure from independent audit.

The 14-field schemas of all eight hypotheses passed format validation. Passing a format check does not establish truth or novelty.

13. M8 — The final ontology verdict

M8 did not take a simple vote over the preceding results. It asked whether any residual remains after every typed baseline.

Candidate basisObserved resultM8 interpretation
Physical responseRicher than the passive quotientAbsorbed by B3/B4 testing and routing
Emergent persistenceNo stable classZ1 v2 collapsed
External relationPositive geometry cross-fitNegative behaviour; not an internal carrier
Carrier registryRich toolkitSurvivor count 0/7
FormationObject-relative definition availableNo universal target
UARTwo subordinate results surviveComposite monomorphism retracted

Therefore,

M8 verdict=NO\boxed{\text{M8 verdict}=\textbf{NO}}

was adopted. This NO is a conclusion within the finite registry and the declared baseline scope.

14. Canon 24's post-verdict theorem

Even after M8, the observer-relative role problem could still be characterised exactly. For role-conditioned transcript laws PL,PIP_L,P_I, the equal-prior Bayes risk is

R=1PLPITV2R^*=\frac{1-\lVert P_L-P_I\rVert_{\mathrm{TV}}}{2}

Adaptive distinguishability is the supremum of total variation over admissible policies. Separately, data processing contracts total variation when an internal transcript is a garbling of an external transcript. Under a fixed infinite policy, finite-prefix error tends to zero if and only if the full path laws are mutually singular. These distinct results did not reopen M8: they establish limited-observer identifiability boundaries, not a new ontology.

15. Lessons from reproduction and audit

  • Saved-output computations matched 13/13 stdout hashes in the current environment.
  • The U0-v1 validator passed 12/12 checks covering preregistration, raw shape, split, gate, verdict, and deterministic replay.
  • The finite verifier for the observer theorem passed 8/8 checks, including distribution pairs, garbling, and adaptive policies.
  • Code reproduction does not replace the full input-provenance and external-access contract.
  • Independent audit overturned an author-side self-audit finding of “not a result defect” and identified a conclusion-level defect.

The current empirical discipline therefore treats author-side closure, independent audit, and LEDGER promotion as distinct stages.

Final empirical conclusion

Strong alignment phenomena and exact boundary theorems exist. Yet on the current data, the measured QQ gain over B3/B4 was nonzero but below the preregistered nonabsorption threshold; no neural-specific ULR candidate therefore retained preregistered meaningful incremental value beyond B*. The empirical programme did not fail: its present achievement is that it maintained null tests demanding enough to overturn its preferred ontology and reach a NO verdict.