# exp013_aug — augmenting pretrained models with the aleph Sequel to [exp012](../exp012_ar/) ([article](https://huggingface.co/blog/AbstractPhil/aleph-autoregressive-differentiation-ft1)). exp012 certified the address bottleneck on a self-trained byte trunk; exp013 asks what the same structures do to FROZEN pretrained models: CLIP-L (both of the last two layers), BERT-base, and GPT-2 124M. All runs single-RTX-4090, pure Adam wd=0, substrates frozen throughout; full ledger in `results/results.json` (`build_results.py` re-asserts every headline below). ## Headline: GPT-2 augmentation (Track C) | arm | ppl (s0 / s1) | mean | trainable | |---|---|---|---| | frozen GPT-2 | 38.648 | 38.65 | 0 | | **aleph relay adapters** | 26.492 / 26.567 | **26.53** | 1.18M (<1%) | | MLP adapters (param-matched, zero-init) | 27.459 / 27.053 | 27.26 | 1.19M | Near-zero-gated multi-slot M̂ relays after every block beat the capacity-matched ablation 2/2 seeds with no overlap — and the gates expose the mechanism: **every aleph gate grew ~3× above its 0.047 init while every MLP gate shrank below it** (asserted per-gate in the ledger). The frozen trunk opens to the address read and closes on the plain transform. Stability asymmetry, same direction: the MLP ablation diverged at lr 1e-3 until given standard output-zero-init; the sphere-normalized relay needed no aid (see `instrument_notes` — the invalid first ablation attempt is recorded, not hidden). Relay codebook drift is depth-graded (0.28 shallow → 0.54 deep): cultivation concentrates toward the prediction gradient inside a pretrained transformer, replicating exp012's depth law in a new regime. ## Track A: token-AR from frozen CLIP-L (both last-two layers) val CE (nats, 49408-vocab; transforms matched ~197K; 2 seeds): | layer | linear | mlp | aleph | sign | |---|---|---|---|---| | final | 5.326 | 5.298 | 5.383 | 5.368 | | penultimate | 5.404 | **5.245** | 5.339 | 5.330 | - **The bottleneck prior is substrate-scoped.** The plain MLP wins frozen-substrate token-AR in every cell; the aleph pays ~+0.09 CE. exp012's 7/7 win required the trunk to co-train with the head; on a frozen semantic substrate the address bottleneck taxes instead. (Preregistered fork; decisive either way.) - **The layer law.** Penultimate + nonlinear is the best cell and penultimate + linear the worst, both seeds, all heads: the penultimate layer is *richer but nonlinearly coded*, while the final layer is more linearly readable — the measured shape of the diffusion-community preference for the penultimate layer. - **Sign > soft aleph in 12/12 pretrained-substrate cells** (8 here + 4 in B1), inverting exp012's long-budget ordering: on frozen substrates the discrete sign-code read is consistently the better aleph consumption. - Cultivation peaks on the penultimate layer (binding fraction up to 0.44) while its task score trails — capacity and cultivation decouple here too. ## Track B1: the joint-failure probe (spelling-AR from pooled embeddings) Decode a word's characters from ONLY a head's read of its frozen pooled embedding (10k wikitext words, exact-match, matched transforms). The gate — traditional heads must actually fail — passed everywhere: | substrate | linear | mlp | aleph | sign | |---|---|---|---|---| | clip_final | .114 | .129 | .141 | **.143** | | clip_penult | .069 | .042 | .044 | .051 | | bert_cls | .009 | .005 | .008 | .011 | | bert_mean | .007 | .004 | .012 | **.013** | The addressed heads extract measurably more of the surface residue that exists (+11% relative on clip_final; 2–3× BERT's floor) but conjure nothing where the embedding genuinely lacks characters — the honest scope of "structure solves head problems" is *extracts more of what is present*: the address is a lookup key, not a compressor. Side-reading: CLIP's final layer retains far more character residue than the penultimate layer or any BERT pooling. ## Specimens `specimens/` holds all 18 checkpoints (Track A/B1 heads + both GPT-2 relay stacks — 40 aleph codebooks total). Every book reads **PROJECTIVE-MOSTLY** under antipodal collapse (`results/specimen_reads.json`): near-uniform on RP³ with emergent antipodal pairs — the projective-codebook law now covers books cultivated *inside a frozen pretrained trunk* and on pooled semantic embeddings. ## Files - `exp013_augmentation_bed.py` — the bed (all tracks/arms, caching, vitals, specimen saving). - `geolip_vitals.py` / `ar_differentiation_bed.py` / `read_codebook.py` — this package's own copies of the shared harness (readout diagnostics; the AlephAddress + MslRelay attention mechanisms; the projective specimen reader). Identical to the `exp012_ar/` copies — the package is standalone. - `repro.py` — loads the code files from this folder and runs them. - `build_results.py` → `results/results.json` — full ledger, self-asserting. - `results/specimen_reads.json` — projective reads of all specimens. ## Reproduce (from inside this folder) ```bash pip install torch --index-url https://download.pytorch.org/whl/cu128 pip install transformers tokenizers pyarrow huggingface_hub python repro.py # CPU smoke: parse + shapes, no downloads python repro.py --track b1 # the gate + baselines (GPU) python repro.py --track a # CLIP-L dual-layer token-AR (GPU) python repro.py --track c # GPT-2 relay augmentation (GPU) python read_codebook.py specimens/ # projective reads of shipped specimens ``` Data + model caches land in `./data` (override with the `GEOLIP_DATA` env var); the first verdict run downloads CLIP-L / BERT / GPT-2 from the Hub. In Colab, paste `geolip_vitals.py`, `ar_differentiation_bed.py`, then the bed into cells in that order — same files, no path setup. License: MIT · AbstractPhil + Claude Fable 5 · July 10, 2026