# Kuku Yalanji v23 Attested-Narrative Adaptation Preregistration

Status: **frozen before RunPod training or v23 inference**  
Date: 2026-07-14 UTC  
Experiment: `v23.0-attested-narrative-adaptation`

## Research question

Does a low-learning-rate second-stage LoRA over the retained v21.2 model improve English-to-Kuku
Yalanji translation on source-attested, speaker-disjoint natural narrative, without materially damaging
everyday usage, dictionary coverage, the elder diagnostic, or the existing synthetic competence?

The treatment is not another exposure-length point on the v21.2 synthetic/Bible trajectory. It adds a
manually reconciled native-speaker narrative corpus and uses restrained retention replay. The Bible has
**zero training rows, zero seed-selection weight, and zero promotion-objective weight**. One Bible set is
run only as a loose catastrophic-forgetting alarm.

## Evidential hierarchy

1. **Primary sealed test:** Patz Text 3, Ivy Walker, Yalanji, 56 clauses. This is opened only after seed
   selection has been written to `seed_selection.locked.json`.
2. **Development and replication:** Patz Text 36, Bobby Roberts, Nyungkul, 53 clauses. This alone selects
   checkpoints and one of three training seeds.
3. **Culturally important nonblind diagnostic:** 43 rights-cleared elder sentence pairs. These rows have
   been inspected in earlier work and are not represented as a blind test.
4. **Behavioral controls:** 84 held-out database usage examples, the 297-row multi-reference dictionary
   probe, and the frozen tagged/untagged synthetic tests.
5. **Bible:** one 325-row direct-translation set, used only to detect catastrophic forgetting at a
   deliberately broad minus-10 chrF++ margin. It cannot select a seed or compensate for a natural-test
   failure.

No metric is pooled across these strata. In particular, Bible volume cannot dominate model judgment.

## Frozen corpus

The deterministic builder is `training/translation/build_v23_attested_adaptation.py`.

| Split or component | Rows | Role |
|---|---:|---|
| Texts 51 and 12, Charlie Tayley, Nyungkul | 156 unique / 624 after fixed x4 replay | New attested train signal |
| DB usage, direct mode | 356 | Everyday/project-database train signal |
| DB usage, original glossary mode | 356 | v21 task-mode retention |
| Synthetic retention sample | 1,024 | Existing competence retention |
| **Training total** | **2,360** | Six epochs per seed |
| Text 36, Bobby Roberts, Nyungkul | 53 | Validation and seed selection |
| Text 3, Ivy Walker, Yalanji | 56 | Sealed natural test |
| Bible | **0 train rows** | No training role |

Nine Text 51/12 clauses carrying explicit transcription/source uncertainty warnings are excluded from
training. A tenth candidate clause, `51.87` (*Yuwu*, "Yes"), is quarantined because its normalized target
duplicates held-out `3.52`. Nine DB rows and six synthetic rows are quarantined by the frozen exact and
token-overlap rules. DB word IDs are disjoint from the 84-row heldout.

### Why the 795-row XIGT export is not training data in v23

All 795 records are retained in `audit/xigt_exclusion_audit.json`, but none is approved for this run.
The export mixes words, paradigms, phrases, clauses, unresolved target/gloss interleaving, and known
systematic OCR substitutions (`l/J/1`). A mechanical screen labels 557 rows as superficially clean, but
that does not establish correct Kuku orthography or clause-level verification. Calling all 795 records
"attested sentence pairs" would overstate the evidence and risk teaching OCR errors. A later version may
use them only after independent normalization and review.

## Leakage controls

Before training, candidate training rows are compared with the natural validation/test and inherited
heldouts using:

- normalized English exact equality;
- letters-only normalized Kuku surface equality;
- English token Jaccard at least 0.85 when both rows have at least four tokens;
- Kuku token Jaccard at least 0.80 when both rows have at least three tokens;
- DB usage word-ID disjointness.

Every exclusion and reason is additive in `audit/quarantine.json`. No held-out or test file is passed to
`train_nllb_lora.py`.

## Model and training treatment

- Base: exact v21.2 step-4,155 merged model, model SHA-256
  `7f9d0fe325e9e4568e45f13179adb336b93bbd53e83ddab2826e999eba3c76f7`.
- Three fixed seeds: `17`, `42`, `73`.
- One RunPod GPU; each seed is trained and evaluated in the same sealed job.
- LoRA rank 32, alpha 64, dropout 0.05.
- Modules: `q_proj,k_proj,v_proj,out_proj,fc1,fc2`.
- Six epochs; batch 8; gradient accumulation 2; effective batch 16.
- Learning rate `1e-5`; warmup 0.05; weight decay 0.01.
- Evaluation/save every 148 optimizer steps (approximately one epoch).
- Best checkpoint within each seed is selected only by Text 36 validation chrF++.
- Full deterministic trainer mode; deterministic float32 inference; TF32 disabled.
- Fixed guarded decoder: one beam, no-repeat 4-gram, repetition penalty 1.10, length penalty 1.0.
- No decoder search in v23.

## Seed selection lock

All three seeds must produce 53 nonempty Text 36 predictions with no tenfold segment loop and mean token
length ratio in `[0.5, 2.0]`. Among eligible seeds, select maximum Text 36 corpus chrF++; tie-break by mean
sentence chrF++, lower maximum repetition, then lexical seed label. The selection artifact is written
before either baseline or candidate inference on Text 3.

## Promotion gates

A research-candidate promotion requires every gate below:

- all three seeds safety-eligible;
- at least two of three seeds improve baseline Text 36 corpus chrF++;
- selected seed improves Text 36 corpus chrF++ by at least 0.5;
- Text 3 corpus chrF++ delta at least +1.0;
- Text 3 paired mean-sentence chrF++ bootstrap 95% interval lower bound above zero;
- more Text 3 row wins than losses;
- positive corpus chrF++ delta on the Text 3 transcription-unflagged slice;
- DB usage corpus chrF++ delta at least -1.0;
- tagged synthetic delta at least -1.0 and untagged synthetic delta at least -1.5;
- elder corpus chrF++ delta at least -1.0;
- dictionary accepted-headword exact count loses no more than three of 297 rows;
- no empty outputs on any evaluated set;
- candidate tenfold-loop count never exceeds baseline;
- Bible direct delta at least -10.0 and no Bible empty output.

The final Bible condition is a catastrophic guard, not positive evidence. A strong Bible score cannot
rescue failure on Text 3, replication, elder, usage, dictionary, or degeneration controls.

## Statistical interpretation

The paired primary estimand is the v23-minus-v21.2 difference in mean per-row sentence chrF++, with a
50,000-replicate paired row-bootstrap percentile interval. Corpus chrF++ and row win/loss counts are also
reported. Rows within a story are not independent speakers, so row-bootstrap intervals do not establish
population-level generalization. There is one speaker in each natural split and a deliberate
Nyungkul-to-Yalanji transfer step. Promotion means only "best current research candidate"; it does not mean
speaker certification, community endorsement, or unrestricted production readiness.

## Frozen artifacts

- Dataset manifest SHA-256:
  `f82f174e4dfbc114f1c7ba6f6c6e14dc87c99a8f4056202a3de11438615e2058`
- Training JSONL SHA-256:
  `6ae80f6525c2bd3d828823cf50aa7bd87dc3ef9d2bcb0b30c90d0c7d01c6c639`
- Validation JSONL SHA-256:
  `34f2c8f28e631a564657e78ba18155ec4b4e0d561d784d1e4d8841cd3c0bd351`
- Sealed test JSONL SHA-256:
  `6d6c1a81ef8be654a8c6d545a70cd5d82a8cc80a63048f0fe658e9513801e6cd`
- Dataset checksum ledger SHA-256:
  `e1494cddcf8b200577e6ed808a5a19885beef27df892c4d70b8a1d41fcf61b77`
- Runner SHA-256:
  `f3847e24e54e9f717d9199102c11b409ea1c377c7ecc834cfb0589af9a841643`
- Promotion verifier SHA-256:
  `d89194549ce3fc2d8e8ea4c42e1ea0d39a2e1f4809cd2e69bb388ea99ae6454b`

The RunPod input ledger will additionally freeze every uploaded code and model file. Any post-freeze code
change requires a new ledger and an explicit amendment before launch.
