loss-manifest / article_loss_manifest.md
AbstractPhil's picture
L-019 retracted (1 dagger) - Expert Soup composite retired with its falsified source system
7adf604 verified
|
Raw History Blame Contribute Delete
72.9 kB

The Loss Manifest: A Field History of Objective Functions, and What a Machine Can Actually Be Asked to Compute

One hundred and fifty-five objective functions, regularizers, gauges, and prohibitions from a multi-year geometric deep learning program — each rated, each with its receipts, failures shipped alongside successes. Read it as a blueprint of a long experiment in computable mathematics: years of drilling machines to find out which structures gradient descent can cultivate in a reasonable amount of time, and which it cannot.


Why keep a manifest of losses

Every result in a deep learning paper is downstream of a decision the paper rarely examines: what, exactly, was the machine asked to compute? The loss is the entire interface between mathematics and learning. It decides what is expressible, what is reachable, and what silently cannot happen no matter how long you train.

This program has been probing that interface for years, across byte-level autoregressive models, CLIP-family alignment systems, diffusion U-Nets and DiTs, vision classifiers, and adapter mixtures on frozen language-model trunks. The through-line was never any single task. It was a single question asked over and over in different geometries: is this structure computable by differential generation — can the gradient path itself cultivate it — or does it only look computable on paper?

Mathematics is generous; optimization is not. A structure can be perfectly well-defined, provably expressive, and still be unreachable in practice — because its gradient homogenizes, because its partition function couples every axis to every other, because its measure collapses without a normalization term, because its worst case is irreducible noise. If everything well-defined were also cheaply computable, this manifest would be short and boring. It is neither, because the universe of trainable mathematics is much smaller than the universe of mathematics, and finding its boundary is empirical work.

So we kept the ledger. All of it: the objectives that carried entire product lines, the regularizers that turned out to be gradient-free the whole time, the elegant formulations that collapsed on contact, the retractions, the prohibitions with their evidence attached. The result is a manifest of successes and failures for a large family of structures — which we would argue is more useful experimental substrate than another benchmark table, precisely because the failures are load-bearing: each one marks a place where the computability boundary was located by direct contact.

Everything below is traceable. The campaign evidence lives in public repositories (the Qwen3.5 adapter line, the Qwen2.5 predecessor line, the diffusion adapter line, the amoe-lora framework, the classification line, geolip-svae, geofractal, the differentiation line), and the story so far is told in three prior field reports: part 1, part 2, and part 3.


The structural finding: three primitives wide, eleven formats deep

The manifest began with a census: three independent sweeps over the program's record — its canon of verdicts, its history layer, and its full code tree, looking for every term that ever received a backward pass. The census found something we did not expect and now consider the organizing fact of the whole record:

Until this week's campaign deliberately built its challengers, only three differencing primitives had ever been back-propagated in this program: cross-entropy, squared error, and KL divergence (plus float64 determinants on the gauge side, and one arm-gated exception). No margin loss, no triplet, no hinge, no load-balancing auxiliary existed anywhere in the historical tree. Everything that looked like a distinct objective was one of those three primitives under a different accumulation format: InfoNCE is cross-entropy accumulated over an N×N similarity grid; structural blob supervision is squared error, dose-coupled and routed to one noise band; frequency-role objectives are squared error under a cosine crossfade. (The campaign then widened the space on purpose — a fourth, Bregman-class primitive and a sparse partition were built precisely to test the boundary, and the margin family was retested under controls. Their verdicts are in era six and in the roster.)

So the program's loss surface is three primitives wide and eleven accumulation formats deep, and nearly every discovery in this manifest lives on the second axis. Two receipts make the point sharply:

  • The worst training collapse in the byte-level line (a coefficients-to-logits head at a single hard temperature: 5.67 bits per byte, address usage perplexity 1.88 of 64 — two winners starve sixty-two axes) was cured by changing accumulation only: split the read into parallel slots and the same cross-entropy lands at 2.47. Identical primitive, different aggregation, night and day.
  • Chunked cross-entropy — sum per chunk, divide once by the global token count — is mathematically identical to plain cross-entropy and operationally a five-fold memory law: on one recorded configuration, 22.8 GB of dedicated VRAM plus 42.8 GB silently spilled to shared memory became 8.8 GB at 1.03 seconds per step. A law of practice that lives entirely in the reduction schedule.

The eleven formats range from the uniform mean (honestly dominant: 84 of 155 entries — plain means are the program's default, and its founding objective lives there) through chunk-renormalized, per-sample-then-weighted, band-crossfaded, masked-denominator, dose-coupled, paired-difference, grid-pairwise, and float64-accumulated forms, down to two cells empty of working objectives. One (raw sum, no denominator) held nothing for the program's entire history — the scale rides on batch and sequence length, so learning rates stop transferring — until this campaign trained its first member, a worst-position accumulation, and refuted it on schedule. The other holds only prohibitions, by statute, and is the single most informative cell in the grid: accumulation that carries state across steps — EMA codebooks, commitment counters, k-means centroids, the entire VQ-VAE bookkeeping family (van den Oord et al., 2017) — contains exactly two entries, and both are prohibitions. Not one working objective in the program's history has ever needed it. The empirical warrant: the program's learned codebooks stay 125+ of 128 axes alive with the diversity weight set to zero. Where the standard literature reaches for a balancing term, this record says the geometry, correctly constructed, balances itself.


How the ratings work, and what a 10 means

Every entry carries a 1–10 rating, and the rating deliberately does not answer "how big is the effect." It answers: how much would we stake on this term in a new, unseen experiment? Six sub-scores (replication across seeds and substrates; measured potency against the term's own gauge noise; doctrinal fit; cost; instrument risk; and a bonus for findings that are enforced in code, not just prose) feed a fixed lookup table, and then nine hard rules bind, in order. The ones that do the most work:

  • anything never run caps at 3 — a beautiful design does not score on paper;
  • a headline resting on an instrument later shown blind takes a penalty until re-measured;
  • single-seed evidence caps at 6; sub-1% margins cap at 5;
  • any formal retraction floors the entry at 1, unconditionally — retracted entries never compete, they testify;
  • contradictory records get a range, never an average; and every sub-score digit must carry a citation or the entry caps as if unrun.

Three calibration pairs prove the rubric measures what it claims:

  1. The same objective rates 8 and 4 in adjacent rows — frequency-band role objectives judged by a role-aligned gauge versus the same objective judged by an aggregate error that was later shown blind to band structure. The only difference is the instrument. That pair is the price of gauge blindness, made explicit.
  2. The same operator rates 6 and 1 — Procrustes alignment as a mild regularizer beside a real training force (it tightens geometric regularity) versus Procrustes as the training force itself (retrieval 0.000, stuck for thirty epochs). Placement decides load-bearingness.
  3. The most potent term in the census rates 2. InfoNCE (van den Oord et al., 2018) is, by measured effect, the strongest alignment force in the record — swap it in and retrieval goes to 0.999. It is also the loudest gradient in the program: representation banks learn it instead of the signal you wanted, and it is banned from address pathways outright. Potency and trustworthiness are different axes. The plainest entry in the record — mean-squared reconstruction driven to bitwise exactness — rates 10.

A short history, told as computability results

Era one: reconstruction, and the discovery that codes are free

The founding result of the program is that reconstruction pressure alone — squared error, uniform mean, nothing else — will cultivate discrete structure that most of the literature builds special machinery to obtain. A small spherical autoencoder driven by plain MSE converged sixteen noise types simultaneously and ended at bitwise-exact text reconstruction; its learned codebook converged to a sign code — rows equal to ± reference vectors at floating-point precision — with no vector-quantization loss, no commitment term, no EMA, no straight-through trick at the objective level. Read the signs, not the probabilities.

That pair of entries sits at the top of the manifest because everything else in the program leans on it: the addressing mechanism at the center of these lines (a signed softmax over oriented half-axes, with sinh in the numerator and cosh in the denominator) receives its only training pressure through reconstruction gradients. In computability terms: discrete codes are computable by differential generation, cheaply and stably — if the pressure is absolute (match this target) rather than comparative (beat those rivals). That distinction became the program's oldest law, and this week it received its sharpest confirmation yet (see era six).

Era two: geometry as force versus geometry as readout

The program spent a long time learning where geometric structure may be pushed and where it may only be watched. The pivotal discovery is almost embarrassing and we publish it anyway: the historical "volume-regularity loss" — a coefficient-of-variation statistic over Cayley–Menger simplex volumes, believed for months to be shaping representations — turned out to be gradient-free the whole time (a .item() call had severed it from the graph). The geometry it was credited with had emerged on its own. That accident became a law: geometric statistics are readouts, never forces, with exactly one sanctioned exception — a micro-weighted (1e-3, hard ceiling) forward CV term on a specific anchor bank, arm-gated, where bare cross-entropy measurably drifts the regularity band and the micro-force holds it at zero task cost.

The rest of the era's entries chart the same boundary from different sides. Sphere normalization — one line of code, no loss term — ended a family of spectral collapses outright and is the founding case for "geometry is regularization: build it in, don't penalize toward it." Direct gradient descent on pentachoron crystals collapses them to zero; the same crystals held frozen retain full cohesion — placement by construction beats placement by hope. Margin-family heads (SphereFace, CosFace, ArcFace — Liu et al., 2017; Wang et al., 2018; Deng et al., 2019) were explored in a vision-transformer lineage and hit a ceiling attributable to the architecture around them rather than the margins themselves. A "soft hand" objective — a reward that boosts the reconstruction gradient near a geometric target rather than penalizing distance from it — produced the best run of its sweep and one adverse finding worth more than the win: sustained moderate boost teaches the model to optimize for staying in the boost zone. Pressure toward a zone is computable; residence in the zone as a goal corrupts.

Era three: alignment, and the loudest gradient

The multi-system alignment era established two placement laws that the rubric now encodes as calibration pairs. InfoNCE is necessary and sufficient as the alignment force — and catastrophic anywhere near an addressing pathway, because grid-pairwise accumulation makes every off-diagonal element a gradient contributor and the bank learns the loss instead of the sequence signal. Procrustes analysis measures alignability and cannot create it. Knowledge distillation (Hinton et al., 2015) earned a statute with teeth after a genetic-selection experiment: distilling at full weight from near-parity teachers produces inverse evolution — best-of-round degrading monotonically across three generations — because children anchor to teacher level and selection feeds the degradation back. Tamed (weight ≤ 0.25, never on founders, never inside a selection loop without a quality gap), the same primitive is a clean positive, and in its row-routed form it demonstrated genuine dark-knowledge transfer: students matching or exceeding teachers on held-out rule induction while the teachers themselves had merely memorized.

Era four: diffusion, and the conditioning law

The diffusion line contributed the manifest's most transferable single result. Identical auxiliary structural supervision — foreground-masked, low-passed squared error on the recovered clean image, dose-coupled at λ≈1 — pays −5.9%/−3.7% on a rectified-flow trunk and is inert (+0.03%/−1.0%) on an epsilon-prediction trunk, a 125–200× effect ratio at two seeds each. The mechanism is exact: flow recovery of the clean image is linear at every noise level, while epsilon recovery divides by a vanishing signal coefficient precisely in the supervised band (cf. Liu et al., 2022; Lipman et al., 2022). Generalized, this is the conditioning law: an auxiliary term pays only where the supervised quantity is recoverable from the prediction through an exact, well-conditioned map — a computability criterion you can evaluate before spending a GPU-hour, and now the first of two pre-spend gates the program runs on every new loss design.

The same era produced the manifest's best accumulation success and its cleanest inertness. Cosine-crossfade band windows over the noise axis — a partition of unity entering both the forward pass and the loss — manufacture surgically decoupled specialists with no routing loss at all (own-band lesion damage 50–200× cross-band; on a diffusion transformer the edge bands reached cross-damage of exactly 0.0). Meanwhile frequency-reweighted "role" objectives moved nothing (0.05–0.2% margins), and the diagnosis became the second pre-spend gate: their gradients were 99.2–99.7% collinear with the base objective. A reweighting of the same residual is still the same pressure. The payers differ in supervised quantity and mask, not in weight — measured at 0.715 gradient novelty for the paying term against 0.003–0.008 for the inert ones.

Era five: the adapter campaigns, and what cross-entropy actually teaches

Two full adapter campaigns on frozen language-model trunks (a 0.5B and then a 0.8B hybrid vision-language model) supplied the workhorse rows: plain, shift-masked, and chunked cross-entropy; the HuggingFace labels path; and the best-evidenced positive supervision result in the record — derived-steps supervision, where training targets carry worked derivations instead of bare answers. It beat direct-answer supervision on held-out generalization at ceiling (1.00/1.00), replicated across seeds to four decimal places, and produced the campaign's first positive off-domain spillover. The same campaigns minted the laws that guard every later row: the question-space law (training-question space must exceed draws threefold, or the loss teaches memorization — discovered by self-retraction when two "experts" passed an answer-diversity guard while memorizing); the toggle law (all adapters off must be bit-exact to the base model — max logit delta 0.0, enforced in code); and the blend-escape diagnostics that grew into a two-regime dispatch law.

The era also filled the proof set. A "controller" anchor hypothesis was refuted at preregistration bars and its failure mechanism classified. Solo always-on specialist stacks proved mutually destructive at n=48. And the consumption-pattern law localized collapse precisely: it is a property of how coefficients are consumed — a single hard-temperature softmax starves non-winners at any dimensionality — not of the address, and slot-parallel accumulation cures it with the primitive untouched.

Era six: the loss campaign — the boundary, measured directly

Everything above set up the question this week finally asked head-on: is cross-entropy itself the right thing to ask a machine to compute, and if not where, exactly, does it fail? Five results — every trained one preregistered at three seeds on a certified byte-level bed:

1. Cross-entropy's failure mode is its partition function, and the failure is graded. The Hessian of CE in logit space has an exact null direction and a spectrum that collapses precisely as the model commits — the same shape as the conditioning law's vanishing coefficient, now inside the loss itself. We built the substrate-native alternative: since CE is the Bregman divergence of log-sum-exp (Bregman, 1967; Banerjee et al., 2005), and the addressing mechanism's own potential is a sum of hyperbolic cosines, the matching loss is the Bregman divergence of Σcosh — curvature bounded below by one everywhere, no null direction, no partition function, antipodally symmetric by construction.

2. In open field, cross-entropy won — as preregistered bars, 3/3 seeds. The cosh-Bregman code loss lost to CE, lost to CE-through-a-frozen-readout, and lost even to its own no-geometry control. Where CE is healthy, nothing we built beats it, and we publish that plainly.

3. Where CE's coupling is the disease, decoupling is the cure — and it's a dose-response. On the certified collapse configuration, with parameters and compute identical and only the loss swapped: full coupling (CE) leaves address usage at 1.85 of 64 axes with saturation 0.9997; a partially coupled partition (sparsemax — Martins & Astudillo, 2016) decompresses to 23.5 of 64; zero coupling (the cosh-Bregman form) reaches 60.9 of 64, with decoded accuracy climbing 0.11 → 0.417 → 0.456. Monotone on every gauge, three seeds per point. A two-year-old diagnosis blaming the geometry was amended: the collapse follows the loss.

4. The deviant sweep calibrated the instruments as much as the losses. A roster of strange forms — focal (Lin et al., 2017), label smoothing (Szegedy et al., 2016; Müller et al., 2019), a confidence penalty (Pereyra et al., 2017), worst-position and geometric-mean accumulations, an anti-curriculum — was gate-measured and then trained without exception. Nothing beat CE. The gate-refused confidence penalty landed closest to CE of all arms, validating the gate's refusal mode against training reality; the two highest-novelty accumulations failed exactly as flagged (one chases irreducible entropy — the worst positions of natural text are not computable structure, they are noise; the other starves the hard positions and the distribution never forms). The protocol finding: novelty is state-dependent — commitment-dependent forms are invisible to gates run at initialization and must be gated at a trained state as well. And novelty is necessary, never sufficient: it screens out inertness; only the bed decides benefit.

5. The oldest law held its hardest test. The program's one previous attempt at replacing cross-entropy entirely — a pure geometric basin loss set (attraction, repulsion, margin, range) from an earlier classification line — was recovered verbatim from its original source and retested under full controls, including CE run on the identical score head. The full set was refuted decisively (−68% relative accuracy — its historical −12% showing was flattered by its original substrate). But the arm that dropped the two roster-comparative terms and kept only the absolute ones more than doubled the full set (0.349 vs 0.157 accuracy, 3/3 seeds). Absolute beats relative, confirmed inside the CE-replacement family itself: the comparative terms are the poison. The bed's loss landscape now shows two clean clusters — absolute-target objectives at 3.75–4.13 bits per byte, partition-coupled cross-entropy at 2.48–2.61 — with the space between them mapped by the coupling dial.


What is computable, then?

Collapsing the manifest to its computability verdicts:

Computable by differential generation, certified here: discrete sign codes via pure reconstruction; surgical specialist structure via positional crossfade windows (no router, no balancing term); structural supervision wherever the supervised quantity is linearly recoverable; derived-steps reasoning supervision; decoupled per-axis code supervision at coupling-driven collapse sites; and near-uniform codebook aliveness with no diversity pressure at all.

Not computable in reasonable time, located by direct contact: gradient-learned alphabets (they collapse; fitted-frozen ones differentiate); pentachora under direct descent; selection events in the compute path (comparative selectors homogenize their own gradient); hierarchy imposed in class space (below-chance, not merely taxed); worst-position training on natural text (the worst positions are irreducible); target-zone residence as an objective; and balancing-by-loss of things that balance themselves.

Conditional, with the condition now measurable: every auxiliary term, via the conditioning gate (is the supervised quantity recoverable through a well-conditioned map?) and the collinearity gate (is the added pressure a genuinely different direction, novelty ≥ 0.3, or a reweighting in costume?). Both gates are calibrated against known outcomes (refusal fires below 0.05 novelty; the payer class begins near 0.3). One refused a design this week that training then confirmed was inert; the other reproduces, retroactively and exactly, the split that cost the record a wasted arm before the gate existed.

That is the blueprint this manifest offers: not a leaderboard, but a mapped boundary — with instruments for extending it that cost seconds, not GPU-days.

Using the roster

The tables below address every entry in the record: what it is mathematically, what it was for, what became of it, where it can be run today, and why it holds its rank. Ratings answer "how much would we stake on this in a new bed." The daggers mark the proof set — retractions and prohibitions kept as first-class citizens, because each is the evidence for a standing law, and because a manifest of failures with receipts is the part of the record you cannot get from papers that publish only what worked.

The open cells are marked too, and they are the invitation: the untested program-native forms (a hyperbolic-distance objective for a sinh/cosh-native substrate; rotor-decode reconstruction; simplex-closure sequence losses), the gauge-to-force promotions each pinned to the lesson of the force that was never a force, and the coupling dial between the two measured clusters. We will be working from this list in the coming days; it is published so that anyone can.


THE ROSTER — all 155 entries

Rated 10 — certified bedrock (17 entries)

ID R entry the mathematics verdict on record lives at
L-001 10 MSE -> bitwise reconstruction (SVAE H2, 16 noise types) L = mean((dec(z) - x)^2); convergence endpoint = bitwise-exact recon 16 noise types converge simultaneously; bitwise text recon; the two-year survivor geolip-svae
L-002 10 recon gradient through M-hat (the aleph's ONLY codebook pressure) M_hat = sum_k sinh(u_k)A_k / sum_k cosh(u_k), u = cos(x,A)/tau; L = mean((dec(M_hat)-x)^2); codebook grad ONLY… cos .992-.997 hard-mode, 125-126/128 axes alive, ZERO collapse, div_weight=0 amoe-lora
L-044 10 Devil's Staircase alpha-normalization (bit_k = p[RIGHT] + alpha*p[MIDDLE], alpha=0.5) p = softmax(-(y-[.5,1.5,2.5])^2/.25); bit_k = p_R + 0.5*p_M; C = sum bit_k 2^-k WITHOUT the alpha term the measure COLLAPSES to {0, .333, .667} - the FractalDavid bug campaign loss library
L-051 10 pure Adam, weight_decay = 0 (the anti-regularizer law) Adam(params, lr, weight_decay=0.0) - the ONLY constructor Adam+gates .731 vs AdamW(3e-4, wd .01) .667 - 'weight decay is uniform damping that destroys the geometric harmonic' amoe-lora
L-052 10 zero-init output heads (WEIGHT and bias) - the inertness contract zeros_(head.weight); zeros_(head.bias); gates = -3.0 makes the toggle law bit-exact (max|dlogit| = 0.0); the bias leak alone is a standing +0.5 ppl offset amoe-lora
L-061 10 question-space guard (training-question space >= 3x draws) assert |question_space| >= 3*draws; train-eval overlap <= .05 caught TWO memorized experts that had PASSED the answer-diversity guard (spaces 480 and 248 vs 800 draws) amoe-lora
L-065 10 band crossfade windows as STRUCTURAL positional gating ramp(x)=.5-.5cos(pi*(clamp(x/XF,-1,1)+1)/2); low=1-up1; mid=up1(1-up2); high=up1*up2; edges(.35,.75) XF=.06 own-band damage 50-200x cross-band, 3/3 both seeds - specialists manufactured with NO routing loss amoe-lora
L-071 10 CV as a READOUT (never a force) CV = std(V)/mean(V), V = CM 4-volumes over 200 random 5-subsets, fp64 - READOUT the historical CV 'loss' was GRADIENT-FREE all along - .item() stripped the graph campaign loss library
L-073 10 bpb (bits per byte) - the AR line's verdict currency bpb = mean CE / ln(2) per byte certified band 2.469-2.499; addr_msl64 beats the unrestricted head 7/7 across seeds and budgets campaign loss library
L-074 10 perplexity tax ladder (wikitext ppl delta, one shared gauge) tax = exp(mean CE_512)|adapted - exp(mean CE_512)|frozen on wikitext one always-on stack +9.23/+9.87 · monolith +3.66 · 5-anchor collective +11.0/+12.6 · UNGATED +91.6 Qwen3.5 line
L-075 10 token-F1 (caption distribution-match delta) F1 = 2PR/(P+R) over token multisets vs GT captions 0.408 -> 0.706/0.704 (+0.30, |s0-s1| = 0.0019); the hub checkpoint reproduces 0.706 EXACTLY Qwen3.5 line
L-080 10 toggle law - all anchors off is BIT-EXACT to the base model assert torch.equal(logits_all_off, logits_base) max|dlogit| = 0.0 exactly at 0.8B on a hybrid DeltaNet/full-attention trunk; library-enforced amoe-lora
L-083 10 usage perplexity / axis aliveness (read-only) usage = mean oriented-softmax row; ppl = exp(H(usage)); alive = usage > eps/2K 125+/128 axes alive WITHOUT regularization - the standing refutation of load-balancing auxiliaries campaign loss library
L-085 10 blend-escape ratio (threshold 1.5) and damping ratio (target >= 3x) ratio = mean|delta|_domain / mean|delta|_neutral; escape <= 1.5; damped >= 3.0 specialists damped 5-11x but caption ESCAPES undamped at 0.1004 - the corollary that became the regime law amoe-lora
L-089 10 CV@1000-batches early screen + the 3-tier filter CV at step 1000 -> band {<.30 LOW / .35-.50 MID / >.80 HIGH} + stability + freeze-survival CV at 1000 batches PREDICTS the final band; turnaround ~2h -> ~7 min per config campaign loss library
L-092 10 resolution-invariance flatness (the debugging canary) var(recon MSE) across patch grids 81..4096 - flatness IS the pass 4.5% MSE variance from 81 to 4096 patches; ~1% across a 36-config sweep - ANY shift means an upstream break record only
L-095 10 peak_mem + s/step (the WDDM sysmem-spill tell) torch.cuda.max_memory_allocated + s/step at an early step (WDDM spill tell) the tell is ~100W/450W at '100% util' with no step prints - 42.8GB observed spilled to shared memory Qwen3.5 line

Rated 9 — replicated and load-bearing (25 entries)

ID R entry the mathematics verdict on record lives at
L-004 9 chunked masked CE (512-token slices, sum-then-renormalize) L = sum_chunks CE_sum(h[i:i+512]) / n_live_tokens (ONE global denominator) 22.8GB dedicated + 42.8GB SILENTLY SHARED -> 8.8GB peak @ 1.03 s/step Qwen3.5 line
L-010 9 flow v-MSE (rectified flow, SHIFT-warped sigma) s = warp(u; shift=2.5); x_t = (1-s)x0 + s*eps; L = mse(pred, eps - x0); x0 = x_t - s*v EXACT/LINEAR x0 = x_t - sigma*v is EXACT and LINEAR at every sigma - asserted, not assumed amoe-lora
L-012 9 addr_msl slot-parallel read (P parallel D=4 slots, shared K=64) feats = concat_p M_hat^(p)(slots); logits = W feats; CE. P=4/16/32/64 dose THE ACCUMULATION CURE: 5.6650 (collapsed) -> 2.47 with the primitive held FIXED campaign loss library
L-015 9 derived-steps expert supervision (stepwise-CoT target vs direct target) shift-CE(-100) on stepwise-CoT target sequences vs direct-answer targets +0.79 vs direct +0.63; held-out ceiling 1.00/1.00; seeds matched to 4 decimals (+0.7917 / +0.7916) Qwen3.5 line
L-016 9 blob-LP-x0 structural supervision on FLOW (lambda ~ 1) L = mean_B[ mse_vec + lam*w_HIGH(s01)*blob_lp ]; blob_lp = sum(blob*(LP(x0h)-LP(x0))^2)/(sum(blob)*C); x0h = x… -5.9% / -3.7% two seeds on flow vs +0.03% / -1.0% on eps: a ~125-200x effect ratio amoe-lora
L-041 9 sphere normalization (M = F.normalize(M); ||M||_F^2 = V pins sum sigma^2) M = F.normalize(M, dim=-1) (||M||_F^2 = V pins sum sigma^2) - ONE line, not a loss zero collapses in 400 epochs; V=1024 went from 48 s/ep crashing to 2.0 s/ep stable record only
L-042 9 gradient equalization across heterogeneous geometric towers per tower: g <- g * target/||g|| (equal gradient norms; outputs stay free) without it spreads hit 20 ORDERS of magnitude (fibonacci dead at 2.25e-21 under helix) geofractal
L-043 9 bounded multiplicative alpha (S*(1 + alpha*tanh), alpha <= 0.2, init .024) Sp = S * (1 + a*tanh(f)), a <= 0.2, init .024 - modulate never inject unbounded alpha POISONS the spectrum; bounded modulation costs 2,272 of 16.9M params (0.013%) record only
L-050 9 gradient clipping discipline (0.5 on cross-attn ONLY; NEVER inside an LBFGS closure) clip_grad_norm .5 on cross-attn ONLY; NEVER inside an LBFGS closure unclipped LBFGS closure DIVERGED to G-MSE 7.4e26; safety is line_search_fn='strong_wolfe' record only
L-068 9 paired (row, noise, t) triples - the variance-killing accumulation acc = mean_fp64(res_arm(row,noise,t) - res_ref(row,noise,t)), triples FIXED per row the noise-pair floor is ~0.988 - without pairing the effects this program measures are invisible amoe-lora
L-076 9 precision + invented-attribute rate (the hallucination decomposition) precision = |pred inter GT|/|pred|; invented = |pred minus GT_vocab|/|pred| precision 0.356 -> 0.694/0.705 and invented-attribute rate 0.200 -> 0.136/0.101, BOTH seeds campaign loss library
L-077 9 register probe (sign-code inter-minus-intra Hamming separation) sep_L = mean_ij inter-register Ham(code_i,code_j) - mean intra (diagonal KEPT, +4% bias, comparability) THE PREDICTOR of the two-regime law: registers ~0.2-0.3 blend, domains ~0.35-0.5 specialize diffusion line
L-078 9 sign_fidelity (Spearman of code-Hamming vs true angular distance) Spearman(Hamming(c_i,c_j), arccos|<a_i,a_j>|) over random pairs PROMOTED: separates inheritance from lottery where bpb CANNOT - successors lock at .9555-.9558, spread < .001 campaign loss library
L-081 9 band-lesion surgical test (own vs cross damage) ratio = damage(own band lesion) / damage(cross band lesion) per gauge surgical 3/3 both seeds at 50-200x; on a DiT edge bands hit cross-damage EXACTLY 0.0 amoe-lora
L-086 9 composition score (the controller prereg gauge) exact-match on two-step composite prompts vs single-step controls the chaining wall: components >= 0.96 solo, composite 0.0 for EVERY config Qwen3.5 line
L-088 9 grad_norm_spread (gradient democracy monitor) orders = log10(max group ||g||) - log10(min); dead = groups with 0 reference failure it exists to catch: 20 orders of magnitude across unequalized towers campaign loss library
L-091 9 spectral gauges: S0/S_D ratio, effective rank, the universal attractor S0/S_D spectral ratio; erank = exp(-sum p ln p), p = sigma/sum sigma critical ratio ~6.5 triggers DISCHARGE; universal attractor S0 ~5.1, erank 15.88 +/- 0.04 across 48+ measurements record only
L-094 9 structured-task validity judges (JSON validity, IoU, pair-order, termination) json.parse validity + IoU(xywh) + pair-order + termination-within-window bbox 0 -> 0.6875 valid (0.894 IoU); the FORMAT TRAMPLING signature: 9/12 truncated_no_json Qwen3.5 line
L-097 9 held-out byte accuracy (rule induction) and variant-format recall (the format lock) held-out byte acc under substitution cipher; variant-format recall teachers memorize at 1.000 train but induce at 0.270/0.245 held-out; memorized content is BOUND to surface form campaign loss library
L-098 9 key-durability gauge (nearest-neighbour symbol Hamming + key drift) NN symbol-Hamming between stored and recomputed keys; match@theta=.25 sign-code keys disagree on ~91% of symbols; match rate at theta=0.25 is 0.000 EVERYWHERE campaign loss library
L-099 9 basin mean_cos (BASIN SET AT INIT) mean cos(book_epoch, book_init) across the bank sweep 192-bank sweep: epoch_1 .8632 / best .8635 / final .8615 - delta 0.0017 BELOW the within-phase std record only
L-100 9 cv_reference_check (fp64 parity against the source of truth) |V_fast - V_geovocab2| / |V| at fp64 == 0 required exact parity (relative 0.0) at fp64 against geovocab2, at ~260x the speed campaign loss library
L-138 9⟂ FAC on the partition-collapse configuration (the P4 loss-swap cell) L-070 on the addr_head collapse configuration the certified addr_head collapse DECOMPRESSES under a loss swap alone, 3/3 seeds: usage ppl 1.0-2.7 -> 60.6-61.1 of 64; decoded acc 0.05-0.20 -> 0.45-… campaign loss library
L-139 9 sparsemax on the collapse configuration (the coupling-axis probe) sparsemax_loss on addr_head logits (K=32, hard tau) - only the loss differs from the certified collapse cell THE DOSE-RESPONSE: usage 1.85 (CE, full coupling) -> 23.5 (sparsemax, partial) -> 60.9 (FAC, none); win|cos| .9997 -> .562 -> .132; acc .11 -> .417 … campaign loss library
L-152 9 PureGeometric ABSOLUTE-ONLY (attraction + range; comparative terms dropped) L = (1 - s_y)^2 + 0.1*(relu(s-1)^2 + relu(-s)^2) - no other-class terms at all MORE THAN DOUBLES the full set: acc 0.349 vs 0.157, bpb 3.75 vs 7.52, 3/3 seeds - the comparative terms are the poison campaign loss library

Rated 8 — solid, one caveat from bedrock (15 entries)

ID R entry the mathematics verdict on record lives at
L-003 8⚠ plain full-sequence cross-entropy (packed labels) L = mean(-log softmax(W h)[y]) the workhorse; also the documented geometry antagonist - CE drove the Oct '25 geometric collapse Qwen3.5 line
L-005 8 shift-CE with ignore_index=-100 (prefix-masked instruction rows) CE(logits[:,:-1], y[:,1:], ignore_index=-100) the standard instruction-tuning form across the v35 and q25 lines Qwen3.5 line
L-009 8 eps-MSE (epsilon prediction, stock schedule) x_t = sqrt(abar_t)x0 + sqrt(1-abar_t)eps, t~U{0..999}; L = mse(unet(x_t,t,c), eps); CFG drop p=.1 relay -2.5% over frozen, 2 seeds; relay >= matched LoRA 2-for-2 across substrates amoe-lora
L-011 8 sign-code head addr_mslh64 (fully discrete forward, STE backward) M_hard = sign(cos[argmax|cos|])*A[argmax]; forward discrete, backward soft (M_hard + M_soft - sg[M_soft]); C… bpb 2.4711 vs soft 2.4685 - parity certified 3 seeds; a ~2.8% gap opens at 4x budget campaign loss library
L-017 8⟂ InfoNCE as an alignment force (OFF address paths) sym CE over sims = za@zb^T/0.07 with in-batch labels NECESSARY + SUFFICIENT for alignment: swap it in -> R@1 .999 Qwen2.5 line
L-040 8 1e-3 CV bank loss (arm-gated, S^15 bank ONLY, never the aleph codebook) V = sqrt(clamp(-det(CM(A[idx5]))/9216)); L += 1e-3 * std(V)/mean(V); fp64, fixed seed-0 subsets, S15 bank ONLY holds CV .295-.305 at zero-to-positive task cost where bare CE drifts it to .31-.34 campaign loss library
L-048 8⟂ HP/LP band-role objectives [judged by the ROLE-ALIGNED gauge] low = base + .5*mse(HP3(pred),HP3(tgt)); high = base + .5*mse(LP7,..); composed by band windows [role-aligned … multiband beats the matched monolith ~10% on HIGH-band foreground, BOTH seeds amoe-lora
L-059 8 straight-through estimator on the aleph HARD read M_hard + (M_soft - sg[M_soft]) (STE over an ABSOLUTE reconstructive read) forward fully DISCRETE oriented code, backward soft: hosted books hold cos .992-.997, 112-122/128 hard axes, zero collapse amoe-lora
L-066 8 lambda dose coupling (3-point curve on the blob term) L = base + lam * w_route * aux, lam~1 (3-pt dose curve) 0.5 -> -5.9% · 1.0 -> -8.3% (in bound) · 2.0 -> -8.4% (OUT of the 0.5% common-gauge bound) amoe-lora
L-067 8 fp64 gauge accumulation (autocast disabled in the reduction) reduce in float64, autocast off (gauges) fp32 determinants lose up to ~4% on near-degenerate pentachora - 'fp32 det only' now means fp32 MINIMUM campaign loss library
L-072 8 anchor drift -> 0.29154 rad + binding_fraction drift = arccos(<norm(a), norm(a_init)>); binding_frac = mean(|drift-.29154|<=.05) the binding constant recurs across 5 architectures and 3 paradigms - but the drift-based fraction is a STAGE statistic campaign loss library
L-079 8 role-aligned in-bed gauge (HIGH-band foreground-masked LP-x0) HIGH-band foreground-masked LP-x0 error (fp32 judged) PROMOTED: found a ~10% multiband win that EVERY aggregate comparison hid amoe-lora
L-082 8 repeated-key null + matched-vs-mismatched deltas excess = metric(real keys) - metric(SAME key repeated); + matched-vs-mismatched delta the instrument that falsified address-as-key: routing excess 2.5e-06 over the null diffusion line
L-096 8 consensus drift / stationarity gauge drift_g = arccos(<consensus_g, consensus_prev>); stationarity = no acceleration ROBUST for structured configurations (0.003 drift by g2, both seeds) but SEED-DEPENDENT for a lone flat book record only
L-149 8 CE on the cosine-anchor basin head (the geobasin control) CE over logits = cos(normalize(feats), normalize(A_c)) * 10 the head itself costs +0.13 bpb under CE (2.607 vs 2.477 linear, 3 seeds; acc .498 vs .505) - small, so every geometric-arm deficit is THE LOSS, isola… campaign loss library

Rated 6–7 — measured, capped by replication or instrument (30 entries)

ID R entry the mathematics verdict on record lives at
L-006 7 HuggingFace out.loss (VLM labels= path, vision tower fires) model(**batch, labels=y).loss (masked shift-CE inside HF; vision tower fires) required wherever the vision tower must fire - chunking bypasses it Qwen3.5 line
L-023 7 kd_facts (fact rows supervised ONLY by teacher logits, alpha=1.0 legal here) fact rows: KL(teacher) ONLY (CE masked off); clean rows: CE - row-routed channels recall 0.953 vs direct 0.871; held-out RULE induction 0.264/0.279 >= the teacher itself campaign loss library
L-024 7 dual-teacher Procrustes consensus distillation GPA: mean shape after per-teacher Procrustes to consensus (delta<1e-8); student anchors init from it teachers .699/.649 -> student .761 EXCEEDS BOTH, still accelerating at E30 campaign loss library
L-028 7 masked-marginal variant scoring (protein VEP) score(v) = logP(x_i=v · x_masked) - logP(x_i=WT · x_masked) (masked marginal) WT unmasked marginal rho 0.10 -> masked marginal ESSENTIAL; final rho .993 / .309 unseen collaboration (pending release)
L-029 7 GPT-2 frozen-trunk relay objective (dif-e013 Track C) CE; trainable = aleph MslRelay adapters on frozen GPT-2 (<1%) frozen 38.648 -> aleph 26.53 vs param-matched zero-init MLP 27.26; beats matched 2/2 seeds campaign loss library
L-055 7 Cayley orthogonality constraint + Newton-Schulz whitening Q = (I-A)(I+A)^-1, A skew - det=1 by construction Q = (I-A)(I+A)^-1 guarantees pure rotation: det = 1.000 throughout, wins 76/84 unseen assays collaboration (pending release)
L-090 7 void topology beta_2/axis (persistent homology on RP^(D-1)) ripser H2 on d(a,b)=arccos|<a,b>| (RP metric), thresh 20deg; beta2/axis within the D=4 cohort every GEOMETRIC signal collapses while VOIDS rise; beta_2 vs recon MSE |rho| = 0.471 record only
L-013 6 addr_3tau multi-tau stroboscope reads at multiple tau; concat -> logits; CE (stroboscope) 4.2884 no collapse (usage ppl 7.9, 117/128 alive) against addr_d4's 5.3698 campaign loss library
L-014 6 addr_mhat reconstructive read consumed in AR logits = head(M_hat) directly (reconstructive read consumed in AR); CE 5.1300 bpb but the HEALTHIEST cultivation on the bed (ppl 11.0, binding_frac .234) campaign loss library
L-018 6 blueprint composite (InfoNCE 1.0 + Procrustes_SVD 0.3 + |CV-0.20| 0.05) InfoNCE*1.0 + Procrustes_SVD*0.3 + |CV(bank)-0.20|*0.05 BERT-8192 m_acc .927 at CV exactly 0.200; CLIP-ctx576 m_acc .945 record only
L-022 6 logit-KD at alpha <= 0.25 with founder exemption L = CE + a*KL(log_softmax(student), mean_k softmax(teacher_k).detach()), a<=0.25, never founders mlp_kd lineage 2.4106 -> 2.3707 -> 2.3662 -> 2.3594 monotone ascent; replicates at s1 campaign loss library
L-026 6 soft-hand loss (proximity REWARD, not penalty) prox = exp(-(cv-target)^2/2sig^2); L = (1+boost*prox)*mse + pen*(1-prox) V256 D24: MSE 0.034 at 400ep - 37% better than the best unconstrained run (.054) campaign loss library
L-027 6 antipode-conv objective (the address AS the convolution operator) conv := fold(m_hat(unfold(x))); no plain filter, no ReLU; CE on head CIFAR-10 87.23% @ 861,450 params with NO ReLU/GELU anywhere; none -> mag +21.8 classification line
L-034 6 entropy-balanced alignment cultivation (w = .05) w=.05 entropy-balanced alignment (exact form NOT fully recorded); endpoint M = +/-ref EXACT produced the emergent basin M = +/- ref EXACTLY - the sign-code convergence endpoint record only
L-035 6 rectified-flow velocity objective (KSimplex / Form 7 bottleneck) rectified-flow velocity mse + Min-SNR gamma=5 + CM terms (L-045/L-046) loss .1749 beat the 268M skip's .1757; the model routed 88% through the 768 dims diffusion line
L-037 6 denoiser objective (tokendiff iterative image-token denoise) CE on x0 tokens from noise-level-t corrupted tokens, iterative beats identity at every level; t=1.0 gives 0.378 vs 0.002 (189x) Qwen2.5 line
L-038 6 recon_target (absolute MSE to a fixed frozen-trunk projection) L = mse(ea, norm(frozen_h @ fixed_proj)) + mse(eb, ...) (absolute target regression) recall@1 0.264 - real (5x frozen) but HALF of InfoNCE's 0.494 at matched budget Qwen2.5 line
L-045 6 L_CM - Cayley-Menger validity hinge (lambda = .01) L_CM = .01 * relu(eps - vol2(CM)) on first k+1 tokens (validity hinge) CM validity 100% across the lineage table diffusion line
L-046 6 L_vol - volume-spread REWARD (-std(log|vol^2|), lambda = .005) L_vol = -.005 * std(log|vol^2| across layers) (spread REWARD, anti-collapse) fragmented anatomy -> coherent composition; base fully preserved (purely additive) diffusion line
L-047 6⟂ Procrustes_SVD as a REGULARIZER (x 0.3 alongside a real force) L = ||A R* - B||^2, R* = Procrustes(A,B) via SVD - as x0.3 REGULARIZER beside a force tightens CV (.19 vs .25) when it rides alongside InfoNCE campaign loss library
L-049 6 anchor dropout (30%) dropout(anchors, p=.3) during alignment prevents collapse: 508/512 anchors active record only
L-054 6 quaternion composition as a structural regularizer (Hamilton product) q_comp = R (Hamilton) q_expert over 4 FiLM arms GeoQuat 0.916 -> 0.993 over 100 epochs vs best baseline 0.903 collaboration (pending release)
L-056 6 cascade as a regularizer (multi-step MLP instead of a direct dimensional jump) k-step MLP cascade in place of one dimensional jump 9-step 256->64 gives 84.6% vs a direct jump's 29.6%; a 27-step r=.95 cascade EXCEEDS the root record only
L-057 6 Cantor router (soft weights derived FROM triangulation distances) w_route = f(phase-0 triangulation distances), softmax-free, geometry-derived cos .9818 at 8 layers vs relay-alone .6533; geometry IMPROVES with more tokens record only
L-060 6 data-level dampening (sqrt damping alpha=0.5, max_repeats=8, cap 1.25x) n_i_new = min(ceil(norm * n_i^0.5), 8, 1.25*top) (sqrt-damped repeats) NEVER equalize-to-largest: alpha=0 repeats 5 images ~50x/epoch record only
L-064 6 rose loss (role-weighted pentachoron regularization, rose_w = 1e-4) NOT RECORDED (role-weighted pentachoron regularization; rose_w=1e-4, temp .07) 74.87% CIFAR-100 @ 393,216 params vs ~65% zero-shot and ~70-72% linear probe record only
L-084 6 read perplexity + |cos to nearest atom| (the quantizer gauge) read ppl = exp(H(mean read weights)); commitment = |cos(read, nearest atom)| read perplexity 14/64 atoms, |cos to nearest atom| 0.964, 64/64 alive - the representation LIES ON the codebook classification line
L-087 6 adapter_effect_mean - the VACUOUS guard effect = mean|loss_off - loss_on|; report VACUOUS if < eps instead of a ratio returns VACUOUS instead of a false PASS when the stack barely moves the loss amoe-lora
L-093 6 exec judge (guarded subprocess: restricted builtins, length cap, hard timeout, no network) guarded subprocess: restricted builtins, len cap, timeout, no net; exact-match out the write-0.0 floor was verified GENUINE off-pod, not a judge artifact Qwen3.5 line
L-104 6 sign-code Hamming retrieval recall@k under Hamming(code_query, code_bank) 0.359 @1 against the continuous head's 0.494 - ~73% of its power from raw 64-symbol Hamming Qwen2.5 line

Rated 2–5 — conditional, refuted-as-candidates, or unrun (38 entries)

ID R entry the mathematics verdict on record lives at
L-007 5 dispatch-keys-only CE (aligner; adapters frozen as anchors) same CE; trainable set = per-block dispatch key matrices ONLY trainable set is ONLY the per-block key matrices; reference-grade, never seed-replicated amoe-lora
L-019 1† Expert Soup composite (InfoNCE + MSE + BCE + Procrustes + CV + spread) InfoNCE + MSE + BCE + Procrustes + CV + spread (6-term, never ablated) RETRACTED 2026-07-31 with its source system — the headline figures were never independently audited and the system carrying them was falsified (retrieval measured an evaluation-protocol leak) record only
L-020 5⚠ SequenceReconstructor loss: MSE(normed) + (1 - cos) mse(norm(pred), norm(tgt)) + (1 - cos(pred, tgt)) on (B,77,768) CLIP-L ep5 m_acc .957 / s_cos .734; Meridian bigG s_cos PLATEAUS at .425 campaign loss library
L-021 5 Min-SNR gamma=5 weighting + velocity adjustment w = min(SNR,5)/(SNR+1) velocity-adjusted; L = mean(w * mse_vec) part of a working recipe (1 ep, 10k synthetic, ~7 min on an L4); never ablated diffusion line
L-025 5 projective-ICP / GPA consensus operator (lineage-core overwrite) projective ICP: iterate sign-aligned Procrustes on RP; lineage-core overwrite recovers planted truth |cos|=1.000 in 5 iterations; TASK-NEUTRAL on bpb, 2 seeds campaign loss library
L-030 5 val_ce on a frozen semantic substrate (CLIP-L token-AR) CE on frozen CLIP-L token-AR (matched transforms + shared vocab proj) MLP WINS frozen-substrate token-AR (penult 5.245 best); aleph tax ~ +0.09 campaign loss library
L-031 5 pure geometric-basin loss set (coherence/separation/discretization/geometry) SOURCE RECOVERED 2026-07-25: attraction (1-s_y)^2 + 0.5*repulsion sum_{c!=y}(s_c^2) + 0.5*margin relu(max_{c!=… the program's ONE attempted CE replacement - NOW PROPERLY TESTED: refuted on the byte bed (acc 0.157 vs control 0.498, 3 seeds); the absolute-only var… geofractal
L-039 5 contrastive dynamics as a CV-compression force standard contrastive; measured as a CV-compression force 100 clusters / 200 steps at d=128 -> CV .2451 (in band); 10 clusters -> .94 record only
L-053 5 geometric autograd / gradient gating (Form 12 tangential-radial split) g_tang pass; g_radial *= (1-.01); g_collapse *= 1.0 (gradient gating) gradients split tangential (pass) / radial (attenuate) / collapse-direction (attenuate) record only
L-062 5 usage / starvation reweighting (drives DATA sampling, NEVER a loss term) on starvation strike: sampling_weight[starved] *= 2; 3 strikes abort - DATA, never a loss the program's ONLY answer to load balancing: x2 upweight the starved anchor's DATA, 3-strike abort amoe-lora
L-063 5 CFG dropout 0.1 (conditioning zeroed, not empty-prompt) with p=.1: cond <- 0 (zeroed, not empty-prompt) standard in every diffusion bed; never ablated in this program diffusion line
L-070 5⟂ FAC as a PRIMARY sequential objective (cosh-Bregman, replace CE) v = norm(feats)@R^T/t; L = mean(cosh(clamp(v - c_y*mu, -4, 4)) - 1) REFUTED AS PREREGISTERED, 3/3 seeds: fac_lsh 4.13 bpb vs ce 2.48; ce_fixedcode 3.81 beats it; fac_none 3.95 beats it campaign loss library
L-101 5 gate-mean band 0.012-0.03 (advisory, NOT universal) gate_mean = mean sigmoid(g); band [.012,.03] ADVISORY held across 6 architectures and 2 optimizers - then MISSED on a 7th at 0.051-0.061 campaign loss library
L-102 5 aggregate eps-MSE as a band-behaviour gauge mean mse over all sigma - BLIND to band structure (distrusted for bands) DISTRUSTED: moved 0.2% against +0.089 grounding effects in image space, and HID a ~10% multiband win record only
L-124 5 frozen-address conditioning injected beside full text append frozen byte-trigram address beside full text cond real vs deranged -0.0009 beside full text; but ALONE the address steers at +0.0287 diffusion line
L-140 5 sparsemax as a full-bed objective L = -z_y + 0.5*sum_{j in S}(z_j^2 - tau^2) + 0.5 (sparse support S) REFUTED as a general objective: bpb 7.43 / acc 0.331 vs ce 2.4769 / 0.505 (3 seeds) campaign loss library
L-141 5 soft-max / worst-position accumulation (trained) L = T*logsumexp(ce_tok/T) - T*log(N), T=0.5 REFUTED: bpb 4.24 / acc 0.276, 3 seeds - the 0.911-novelty champion chases irreducible entropy exactly as flagged campaign loss library
L-142 5 geometric-mean accumulation (trained) L = mean(log(ce_tok + 1e-3)) REFUTED decisively: bpb 9.03 - the anti-focal starves hard positions and the distribution never forms (3 seeds) campaign loss library
L-143 5 label smoothing eps=.1 (trained on the byte bed) CE to (1-eps) smoothed targets == (1-eps)CE + eps*uniform-KL bpb 2.587 vs ce 2.4769 (+0.11, 3 seeds) - payer-class novelty (0.479), mildly WORSE outcome campaign loss library
L-144 5 focal gamma=2 (trained, live-model weights) L = sum((1-p_y)^2 * ce_tok) / sum((1-p_y)^2), p_y detached from the live model bpb 2.597 (+0.12 vs ce, 3 seeds) - payer-class trained novelty (0.337), mildly worse outcome campaign loss library
L-145 5 anti-curriculum (train only where the frozen reference is confident) L = sum(ce_tok * [pi_ref > .6]) / count, pi_ref from the frozen ce_s0 checkpoint REFUTED as an objective: bpb 6.74 (3 seeds) - abandoning 72% of the distribution buys nothing on the rest campaign loss library
L-146 5 FAC tanh-Hamming link (bounded tails) L = mean(1 - tanh(v) * c) cosh beats tanh 3/3: 4.349 vs 4.1285 (+0.22) - the bounded link loses within the family campaign loss library
L-147 5 FAC Cauchy link (sub-quadratic tails) L = mean(log(1 + (v - c*mu)^2)) cosh beats Cauchy 3/3: 4.360 vs 4.1285 (+0.23) - robust-statistics tails lose within the family campaign loss library
L-148 5 confidence penalty (trained as the GATE-VALIDATION CONTROL) L = CE - 0.1*H(p) CLOSEST TO CE OF ALL DEVIANTS: bpb 2.520 (+0.043, 3 seeds) - the gate's refusal correctly predicted 'CE plus nothing' campaign loss library
L-150 5 PureGeometricLoss, learned anchors (the Oct '25 arm, properly tested) attraction (1-s_y)^2 + 0.5*sum_{c!=y}s_c^2 + 0.5*relu(max_{c!=y}s_c - s_y + .3) + 0.1*range REFUTED on this substrate: acc 0.157 vs control 0.498 (-68% relative, 3 seeds) - far below the historical -12% trade geofractal
L-151 5 PureGeometricLoss, FROZEN anchors (the L-108 cell) same loss; A registered as a frozen buffer learned BEATS frozen by +8 acc points (0.157 vs 0.076, 3 seeds) - the L-108 falsifier FIRED for cosine anchors campaign loss library
L-153 5 GeometricPrototypeLoss (verbatim, own projector) cos(proj(scores), prototypes) pulled/pushed + prototype-diversity term WORST of the family: bpb 8.12, acc 0.008 (3 seeds) - the extra indirection buys total failure geofractal
L-154 5 HierarchicalGeometricLoss on the nibble hierarchy (16x16) coarse (superclass sums to target) + fine + consistency, sigmoid-weighted CATASTROPHIC: acc 0.0003 - below chance (1/256) - hierarchy-in-class-space destroyed fine structure entirely (3 seeds) geofractal
L-155 5 CE + PureGeometric hybrid (0.5/0.5) 0.5*CE(cos*10) + 0.5*PureGeometricLoss(scores) the geometric set POISONS CE rather than riding it: bpb 4.53 vs control 2.61 (+1.9, 3 seeds) - P4 bar (within 0.15) missed by 12x campaign loss library
L-008 4 image-classification CE (CIFAR-10, aleph-dispatched MoE vs dense) CE(logits, y) on CIFAR-10 MoE 58.52% TIES param-matched dense 58.52% exactly; 6x params bought nothing campaign loss library
L-036 4 margin losses ArcFace / CosFace / SphereFace (RoseFace dual-norm) ArcFace cos(th+m) · CosFace cos(th)-m · SphereFace cos(m*th); s=30 m=.30; L1-then-L2 dual-norm the ZANA innovation - and it hit a 60% single-stream ceiling campaign loss library
L-123 4⟂ HP/LP band-role objectives [judged by AGGREGATE eps-MSE] L-048 judged by aggregate eps-MSE 4/4 directional both seeds at 0.05-0.2% margins - 'nearly collinear with the base objective' record only
L-032 3 GBC - 'cross-entropy can be replaced entirely' (roadmap claim) SOURCE RECOVERED 2026-07-25 (GBC head, geofractal/model/experiment_geometric_basin.py:118): compat = triadic (… classification via triadic compatibility, self-similarity, Cantor coherence, hierarchical basin checks geofractal
L-033 3 masked-recon / generative arm (campaign law 2 in its ORIGINAL form) mask patches; L = mse(recon_from_antipode_read(masked), x) (law 2 ORIGINAL form) BUILT, NEVER RUN - predicted to be where the SIGNED read finally beats magnitude classification line
L-058 3 address-agreement bias (BUCKET - making a hard address differentiable) exact softmax within sorted equal-width same-bucket block; codebook grad via address-agreement bias exact softmax within sorted equal-width blocks masked to the same bucket; argmax alone is gradient-dead record only
L-069 3 predictability-weighted accumulation (PWA) w = f(pi_frozen_ref); L = sum(w*ce_tok)/sum(w) DESIGNED 2026-07-25: make the PREDICTABILITY PRINCIPLE a loss geometry instead of a discovered side effect campaign loss library
L-103 2✖ recon cosine as a judge for ADDRESSED systems cos(recon, x) - WRONG instrument for addressed systems (address = lookup key) DISTRUSTED: an address is a LOOKUP KEY, not a compressor - judge drift and crushed CV instead record only
L-113 2✖⟂ InfoNCE into ADDRESS paths a7_grid_infonce INTO an address path BANNED despite R@1 .999 - it is the LOUDEST gradient and the bank learns IT instead of the useful signal campaign loss library

Rated 1 — the proof set (retractions and prohibitions) (30 entries)

ID R entry the mathematics verdict on record lives at
L-105 1✖† VQ / commitment / EMA codebook losses ||sg[z_e] - e||^2 + beta*||z_e - sg[e]||^2 (+ EMA codebook update) THE NAMED PROHIBITION - and unnecessary: the codebook stays 125+/128 alive at div_weight = 0 campaign loss library
L-106 1✖† comparative / relative selectors (argmax anchors, softmax-over-roster, STE one-hots, k-means alphabets) selection event = argmax/softmax-over-roster in the compute path roster-dependent; the gradient HOMOGENIZES - 14x path collapse, width attenuation, BN-on-padding, same disease record only
L-107 1✖† gradient-learned alphabets (CAMPAIGN LAW 3) alphabet learned by task gradient (vs fitted-frozen) fitted-frozen alphabets differentiate (1,594 unique paths); gradient-learned alphabets COLLAPSE (116) record only
L-108 1✖† direct gradient descent on pentachora direct task-gradient descent on pentachoron vertices collapses them to zero - as FROZEN anchors the same crystals retain full cohesion and stay backtrackable record only
L-109 1✖† global average pooling in geometric encoders gap = x.mean(dim=spatial) in a geometric encoder 70% -> 29% collapse, REPLICATED independently in the protein line campaign loss library
L-110 1✖† CV loss as backward injection / above the 1e-3 ceiling CV term injected in backward, or weight > 1e-3 MUST be a forward loss; above ~.001 the CV term dominates CE and trades discrimination for regularity record only
L-111 1†⟂ Procrustes as a training FORCE same as L-047 - AS THE TRAINING FORCE (placement retracted) as a training loss: R@1 = 0.000, P_cos stuck at .094 for THIRTY EPOCHS campaign loss library
L-112 1✖† addr_head - coefficients to logits at a single hard tau logits = W u, u = single-slot coefficients at hard tau (K=32) 5.6650 bpb COLLAPSED: usage ppl 1.88/64, TWO unique winners, win|cos| .9992 campaign loss library
L-114 1† logit-KD at alpha = 1.0 from near-parity teachers prim_kl at alpha=1.0 from near-parity teachers in a selection loop INVERSE EVOLUTION, compounding downward: 2.4301 -> 2.5046 -> 2.5603 campaign loss library
L-115 1✖† blob structural supervision on the EPS objective L-016 with x0h = (x_t - sqrt(1-abar)eps_hat)/sqrt(abar) - divides by vanishing sqrt(abar) +0.03% / -1.0%, two seeds - the x0 recovery divides by a vanishing sqrt(alpha_bar) EXACTLY in the supervised band amoe-lora
L-116 1† MSE-first single-epoch keep-or-kill screening keep-or-kill on 1-epoch MSE rank DEAD: the lowest-MSE config was a HIGH-band false candidate record only
L-117 1† tied M-hat readout (U=M_hat, S=Omega-token, Vt=I) in an AR head logits = tied(M_hat) with U=M_hat, S=Omega, Vt=I +1.0 bpb BOTH seeds and it STARVES the codebook (drift 0.02, binding 0) record only
L-118 1✖† comparative routing on diffusion (state+sigma, raw address, M-hat address-as-key) route experts by frozen text keys (raw/pooled/M-hat-slot) vs repeated-key null FALSIFIED THREE WAYS, 2 seeds: routing excess 2.5e-06 over the repeated-key null; match advantage -0.0 diffusion line
L-119 1† the controller hypothesis (a trainable anchor that orchestrates the others) a trainable anchor trained to orchestrate others (composite prereg >= +.15) prereg required >= +0.15; measured -0.417 / -0.167. The passenger role is an ATTRACTOR Qwen3.5 line
L-120 1† always-on solo specialist stacks solo specialist stack attached always-on MUTUALLY DESTRUCTIVE at n=48: the depth stack drives caption F1 to 0.0014 with termination 0.0 Qwen3.5 line
L-121 1† frozen solo-trained expert collectives under aleph dispatch frozen solo-trained experts composed under dispatch no surgical independence (own-drop 0.04/0.00), NO damping (all five blend-regime, 0.86-1.6), composite 0.0 Qwen3.5 line
L-122 1† organ-only inheritance (projection + book transplanted onto fresh trunks) transplant proj+codebook onto a fresh trunk BELOW random init, 2/2 lineages - sixteen random draws beat component inheritance record only
L-125 1† shuffled-key null null = shuffle keys across rows (measures diversity, not correctness) CONFESSED INSTRUMENT FAILURE: it measures diversity, not correctness - the null scored like the real thing record only
L-126 1† the sequences / baseconv expert gains CE on generated question sets with space < 3x draws SELF-RETRACTED: question space 480 and 248 against 800 training draws per tier = MEMORIZED record only
L-127 1† the exp021 seed-inversion claim for the trainable anchor cross-seed comparison across DIFFERENT instruments RETRACTED WITHIN HOURS: the claim compared DIFFERENT INSTRUMENTS across seeds campaign loss library
L-128 1✖† hierarchical refinement in Cantor space bands nested within bands on a Cantor axis HARMFUL (-10%); parallel ADJACENT NON-OVERLAPPING bands are +3% record only
L-129 1✖† repeated boundary crossing in a measure space re-enter measure space per layer/step KILLS gradients (catastrophic -> random). Enter and exit the measure space ONCE record only
L-130 1✖† the SOFT devil's staircase used as a BAND COORDINATE soft_cantor_ungated(x) used as a band COORDINATE (non-monotone) NEW 2026-07-25: measured NON-MONOTONE - min slope -0.13 to -0.49 at EVERY level count on EVERY grid tested geofractal
L-131 1✖† equalize-to-largest data balancing (alpha = 0) repeat count = ceil(max_bucket / n_i) (alpha=0 equalize-to-largest) repeats the 5-image bucket ~50x per epoch - 'the textbook way to overfit the long tail you were trying to protect' record only
L-132 1† addr_conv - the decorative address (convex re-weighting of a filter bank) conv re-weighted by convex sum a_k=1 over a filter bank (hull-bounded mean) DECORATIVE: a convex sum a_k = 1 is a hull-bounded perturbation of a MEAN; the 1x1 address is CONSTANT on grayscale (variance 4e-16) classification line
L-133 1† deterministic (greedy) decoding in an iterative denoiser argmax decoding in an iterative denoiser collapses to the global mode: diversity 0.0, conditional == shuffled EXACTLY record only
L-134 1✖† load-balancing / auxiliary router losses aux = alpha * N * sum_i f_i * P_i (switch-style balance) BANNED and replaced by architectural equality; ZERO instances exist in the tree amoe-lora
L-135 1† the big-JSON objective CE on the big-JSON composite format FORMALLY DROPPED by operator ruling - too costly; 3-5 task adapters deliver more per GPU hour record only
L-136 1† SVD-rotation transform in the dual-pentachoron head learned SVD rotation transform in the dual-penta head DROPPED for convergence failure; reduced to scale + shift record only
L-137 1✖† single hard-tau coefficient heads at ANY dimension coefficients->logits at ONE hard tau, any dim DEMOTED on the standing registry: collapse, and low-D was falsified as the fix record only

The rated manifest, its machine-readable sidecar, the rubric, the campaign code, and the raw run ledgers accompany this article at the companion repository. Evidence trails for every historical claim live in the linked line repositories; the three prior field reports carry the campaign narratives in full. Written by the program's operator (AbstractPhil) with Claude (Anthropic) as the research engineer of record for the loss campaign.