# Memory, Tokens, and Learning Geometry

**Status:** Technical research questions and candidate models.  
**Edition:** Edited English reconstruction. This document is a narrower technical branch of the [evolving representation-language program](evolving-semantic-representation-language.md). Algebraic examples are not language-model benchmarks.

## 1. What is learned when experience becomes memory?

Can a system learn both the units with which it represents experience and the rule by which it writes experience into useful future state?

Memory storage, memory addressing, interpretation, and learning must be distinguished. A low-rank parameterization specifies a form of change; it does not supply the complete policy for writing or checking that change.

## 2. Parametric memory as a candidate

For a layer ℓ:

    W_(ℓ,t) = W_(ℓ,0) + s_ℓ B_(ℓ,t) A_(ℓ,t)
    A ∈ R^(r×d_in), B ∈ R^(d_out×r)

A changing adapter can alter responses to future inputs. A gate could control how strongly an internal memory branch contributes.

Writing could use optimization on experience or a trained generator that proposes updates. These alternatives move costs to different places. Neither implies unlimited lossless memory at a fixed rank.

Parameter count, stored bytes, optimizer moments, gradients, activations, and write time are separate costs.

## 3. Three kinds of compact unit

| Unit | What must be shared or learned |
| --- | --- |
| Vocabulary identity | A codebook and a trained reader |
| Dynamic latent vector | An encoder and an interpreter |
| Memory address | A store, addressing rules, and version compatibility |

A compact address can activate rich shared knowledge without carrying all of it in the address. The shared store still counts in the resource budget.

## 4. Merging must preserve order and scope

A candidate merger G_φ takes representations and positions of successive units. A simple average need not preserve order, negation, names, or syntactic boundaries.

When predicting a future token, the merger cannot use that token as if it were already observed. Compressing an observed input and generating an unknown continuation are different tasks.

Fewer attention positions do not imply an equal reduction in total runtime. Encoding, local computation, storage, and decoding remain.

## 5. Preserve behavior over a question family

A candidate objective compares future answer distributions under a full history X and a compact representation Z with shared memory M:

    L_keep = E_(q∼Q) KL(p_ref(·|X,q) || p(·|Z,M,q))

The question family Q determines which distinctions matter. A representation adequate for a broad goal may fail an exact-detail question.

A reference model can also be wrong. Independent ground truth should be used where available. The compressor must not see the held-out evaluation questions in advance.

Arbitrary lossless compression into a fixed finite-precision state is not assumed.

## 6. Writing while limiting interference

Consider a linear associative memory M, a new key k, desired value v, and error e=v−Mk.

Let C be the second-moment matrix of older keys and C̃=C+λI, with λ>0. A candidate constrained update solves:

    minimize  ½ tr(ΔM C̃ ΔMᵀ)
    subject to ΔM k = e

For nonzero k:

    u = C̃⁻¹ k / (kᵀ C̃⁻¹ k)
    ΔM = e uᵀ

The constraint makes the new relation exact in this linear problem. A different key k′ changes by e(uᵀk′), exposing possible interference.

This is a limited linear objective, not a guarantee about all older tasks. Conflicting values for overlapping keys may prevent simultaneous exact satisfaction.

## 7. An illustrative two-dimensional comparison

Take C=diag(100,1), k=(1,1), and scalar error e=1. C is already positive definite.

| Update direction | Change at new key | Weighted squared change |
| --- | --- | --- |
| (1/2,1/2) | 1 | 25.25 |
| (1/101,100/101) | 1 | 100/101 |

The ratio is 25.5025 for this chosen metric and example. It is an algebraic illustration of why a metric matters. It is not a measured speedup, a general retention factor, or evidence about a trained language model.

## 8. The cost of the metric

A full d×d second-moment matrix is expensive. Blocks, a smaller key space, or diagonal-plus-low-rank approximations are candidate compromises.

Sequential rank-one writes can increase rank. Compressing back to a fixed budget may discard information important to a later query even when matrix approximation error is small.

Evaluation should therefore include future retrieval and task performance, not only singular values or update norms.

## 9. Three geometries

Representation state, parameter state, and external task state are different spaces. A token's layer trajectory is not the optimization path of the weights. An action graph is not automatically a smooth manifold.

Gradient flow also need not be a geodesic. A metric-dependent steepest-descent direction is local; it does not establish a globally shortest learning path.

The non-uniqueness BA=(BR)(R⁻¹A) for invertible R illustrates why distance between factor coordinates can differ while the effective update is identical.

A behavioral metric may be more meaningful, but its usefulness remains a hypothesis to test.

## 10. Learning a writer

A learned mechanism might propose update keys, values, gates, rank allocation, or step sizes rather than emitting every parameter.

Inputs could include selected representation trajectories, error signals, prior state, and gradients. Using gradients retains their computation cost. A gradient-free generator may move that cost into its training.

The outer objective should evaluate later unseen tasks, retention, correction of obsolete information, and resource use. The writer's confidence is not a sufficient validation signal.

## 11. Communication between different readers

A candidate transfer is:

    z_A = E_A(X)
    z_B = T_(A→B)(z_A)
    y_B = D_B(z_B,q)

Equal dimensions do not establish compatibility. A learned translator or shared interpretation is required.

One latent vector can use fewer attention positions but more bytes than many token identifiers. Shared-memory synchronization, adapter compatibility, and version changes belong in the comparison.

## 12. Continuing state after a write

A parameter change can make a previously computed cache inconsistent with recomputation under the new parameters. Candidate responses include updating at explicit boundaries, refreshing affected state, or using a separate memory branch.

A resumable learner may need optimizer state, concept versions, memory addresses, short context, and task state as well as weights.

A remembered observation about an external service does not establish its present state. Grounding may require a fresh observation.

## 13. A fragment as an operation

Let T_a(S) denote the effect of fragment a on a sufficiently specified state S. Then an operation-preserving merger targets:

    T_G(a,b)(S) ≈ T_b(T_a(S))

Order generally matters. Hierarchical grouping also needs evaluation. A fixed-sized code family is not assumed closed under arbitrary composition.

In a linear toy model:

    T_a = I + U_a V_aᵀ
    T_b = I + U_b V_bᵀ

their composition is:

    T_b T_a = I + U_a V_aᵀ + U_b V_bᵀ + U_b(V_bᵀU_a)V_aᵀ

The interaction term shows why adding two adapters need not preserve sequential effects. The rank of T_bT_a−I is at most r_a+r_b. Recompression can be lossy.

This illustrates algebraic constraints, not a theorem identifying Transformer semantics with linear maps.

## 14. Discriminating experiments

Compare raw context, retrieval, prose compression, latent slots, and parametric memory under the same total budget. Compare writers at the same success on new information before judging retention.

Include changed facts, contradictions, novel names, ordering, negation, exact quantities, and long action chains. Measure one-off and repeated-use costs separately.

The relevant achievement is useful preserved or improved behavior. A lower token count alone does not show a new semantic language.
