Conclusion
There is no single observer-independent YES or NO to the question, “Can learning and inference be
distinguished inside a model?” The exact answer is the following necessary-and-sufficient condition:
An observer can distinguish the current role exactly if and only if the declared role variable is measurable with respect to that observer's information filtration.
If a role induces transcript laws , then the optimal equal-prior misclassification probability is
1. Declare the role and observer first
Learning and inference are not primitive symbols of a probability space. A system boundary must first declare which events it calls or . Role typing may include
The information available to an observer through time is represented by a sub--algebra . This filtration may include activations, inputs, outputs, a clock, observer memory, and the results of permitted interventions.
Location is not the decisive issue.
- An observer located inside the model remains limited if it cannot inspect optimiser state or reset responses.
- An observer outside the model can have the same information filtration if it receives the same latent stream.
- Even access to raw hidden state does not recover the role automatically when the role is not a function of that state.
2. Theorem A — Measurability characterisation
For a probability space , a role , and observer information , the following statements are equivalent:
- an -measurable classifier exists that predicts almost surely;
- is -measurable;
- the posterior is almost surely either 0 or 1.
The minimum risk over all classifiers is
A larger internal network or a longer computation is only a measurable function of the existing -algebra. If the input information does not separate the roles, additional computation cannot create new role information.
3. Theorem B — Exact risk for a fixed transcript
Fix an observer contract and horizon, and let the transcript space be . The optimal binary risk for the two simple hypothesis laws is given exactly by the total-variation identity
Consequently,
The verdict is determined not by different means or weights, but by the entire transcript law that actually reaches the observer.
4. Theorem C — Adaptive experiments
Suppose the observer can use a nonanticipating, role-blind policy that selects actions from the past transcript. Define the separation coefficient at horizon by
The infimal risk over permitted policies and classifiers is
Positive TV is equivalent to performance above chance; finite-time zero error requires singularity under one permitted policy. Even if , the supremum need not be attained, so the existence of a single zero-error policy must be checked separately.
5. Theorem D — Response-fibre factorisation
Let be a set of hidden events and a role map. Identify events that produce the same response profile under every permitted policy and horizon:
A universal observer classifier valid for every event exists if and only if factors through this quotient:
If and conflict within the same response fibre, no amount of observer-side computation can resolve the conflict. Conversely, factorisation establishes in-principle classification by profile; it does not automatically provide a single finite test.
6. Theorem E — Garbling contraction
If an internal transcript is obtained by garbling an external transcript through a Markov kernel , then
Thus, additional computation by a weaker observer cannot recover provenance that has been lost. Neither the spatial claim that “inside always knows more than outside” nor its converse is valid. The declared channel determines the information ordering.
7. Theorem F — Infinite horizon
For one fixed infinite policy, let be the finite-prefix laws. Prefix TV converges to full path-law TV. The optimal finite-prefix error tends to zero if and only if the full path laws are mutually singular:
This is a result for one coherent infinite policy. It must not be confused with a profile that selects a different oracle policy at every horizon.
8. Theorem G — Common quotient kernel
If two processes have the same observable quotient transition and their output kernels do not read the hidden fibre directly, different hidden-fibre updates can remain indistinguishable in the declared adaptive transcript. This result is stated as a sufficient condition. Its assumptions fail if the output reads the fibre directly or if a policy is given the role label in advance.
9. Reset-relative operational specialisation
Let denote the probe law after event , the law after a typed reset , and the baseline law in which the event did not occur. Define
- indicates persistent adaptation relative to this reset contract.
- If but the pre-reset response differs, the event is reset-local inference.
This supplies one operational definition of learning and inference roles. Changing the reset may change the role; it is not a universal definition covering all fast learning and continual adaptation.
10. Strictness of passive and richer ports
In a linear toy model, two systems with the same passive law can separate with TV 1 under a write; reset; read policy.
Merely adding an operation named “reset,” however, is not sufficient. If distinct persistence operators
satisfy
they remain in the same partial response fibre. The identifying power of a richer port is itself interface-relative.
11. Scope of proof and verification
The canonical audit establishes the following scope:
| Requirement | Result |
|---|---|
| Possibility of zero error from current information | Necessary-and-sufficient condition in A |
| Optimal error for a fixed stochastic transcript | Exact TV identity in B |
| Adaptive inputs and interventions | Policy supremum in C |
| Universal classifier over all events | Quotient factorisation in D |
| Direction of information loss between inside and outside | Garbling contraction in E |
| Zero risk under long observation | Path-law singularity in F |
| Elimination of hidden-fibre differences | Explicit sufficient condition in G |
The finite exact verifier passed 8/8 checks, including 225 distribution pairs, 8 deterministic garblings, 27 stochastic kernels, and 64 adaptive quotient policies. This computation does not replace the general measure-theoretic proof; it audits signs, normalisation, and finite counterexample implementations.
What is not proved
- An observer-independent ontological definition of learning and inference
- Identification of the observer filtration in a real Transformer
- Finite-sample estimation rates for a classifier
- That this theorem constitutes a neural-specific ULR object
The Canon 24 increment is therefore not an ontology, but a complete observer-relative characterisation of the identifiability boundary for a declared role.