Kuku Yalanji v23 Attested-Narrative Adaptation Preregistration
Status: frozen before RunPod training or v23
inference
Date: 2026-07-14 UTC
Experiment: v23.0-attested-narrative-adaptation
Research question
Does a low-learning-rate second-stage LoRA over the retained v21.2 model improve English-to-Kuku Yalanji translation on source-attested, speaker-disjoint natural narrative, without materially damaging everyday usage, dictionary coverage, the elder diagnostic, or the existing synthetic competence?
The treatment is not another exposure-length point on the v21.2 synthetic/Bible trajectory. It adds a manually reconciled native-speaker narrative corpus and uses restrained retention replay. The Bible has zero training rows, zero seed-selection weight, and zero promotion-objective weight. One Bible set is run only as a loose catastrophic-forgetting alarm.
Evidential hierarchy
- Primary sealed test: Patz Text 3, Ivy Walker,
Yalanji, 56 clauses. This is opened only after seed selection has been
written to
seed_selection.locked.json. - Development and replication: Patz Text 36, Bobby Roberts, Nyungkul, 53 clauses. This alone selects checkpoints and one of three training seeds.
- Culturally important nonblind diagnostic: 43 rights-cleared elder sentence pairs. These rows have been inspected in earlier work and are not represented as a blind test.
- Behavioral controls: 84 held-out database usage examples, the 297-row multi-reference dictionary probe, and the frozen tagged/untagged synthetic tests.
- Bible: one 325-row direct-translation set, used only to detect catastrophic forgetting at a deliberately broad minus-10 chrF++ margin. It cannot select a seed or compensate for a natural-test failure.
No metric is pooled across these strata. In particular, Bible volume cannot dominate model judgment.
Frozen corpus
The deterministic builder is
training/translation/build_v23_attested_adaptation.py.
| Split or component | Rows | Role |
|---|---|---|
| Texts 51 and 12, Charlie Tayley, Nyungkul | 156 unique / 624 after fixed x4 replay | New attested train signal |
| DB usage, direct mode | 356 | Everyday/project-database train signal |
| DB usage, original glossary mode | 356 | v21 task-mode retention |
| Synthetic retention sample | 1,024 | Existing competence retention |
| Training total | 2,360 | Six epochs per seed |
| Text 36, Bobby Roberts, Nyungkul | 53 | Validation and seed selection |
| Text 3, Ivy Walker, Yalanji | 56 | Sealed natural test |
| Bible | 0 train rows | No training role |
Nine Text 51/12 clauses carrying explicit transcription/source
uncertainty warnings are excluded from training. A tenth candidate
clause, 51.87 (Yuwu, "Yes"), is quarantined
because its normalized target duplicates held-out 3.52.
Nine DB rows and six synthetic rows are quarantined by the frozen exact
and token-overlap rules. DB word IDs are disjoint from the 84-row
heldout.
Why the 795-row XIGT export is not training data in v23
All 795 records are retained in
audit/xigt_exclusion_audit.json, but none is approved for
this run. The export mixes words, paradigms, phrases, clauses,
unresolved target/gloss interleaving, and known systematic OCR
substitutions (l/J/1). A mechanical screen labels 557 rows
as superficially clean, but that does not establish correct Kuku
orthography or clause-level verification. Calling all 795 records
"attested sentence pairs" would overstate the evidence and risk teaching
OCR errors. A later version may use them only after independent
normalization and review.
Leakage controls
Before training, candidate training rows are compared with the natural validation/test and inherited heldouts using:
- normalized English exact equality;
- letters-only normalized Kuku surface equality;
- English token Jaccard at least 0.85 when both rows have at least four tokens;
- Kuku token Jaccard at least 0.80 when both rows have at least three tokens;
- DB usage word-ID disjointness.
Every exclusion and reason is additive in
audit/quarantine.json. No held-out or test file is passed
to train_nllb_lora.py.
Model and training treatment
- Base: exact v21.2 step-4,155 merged model, model SHA-256
7f9d0fe325e9e4568e45f13179adb336b93bbd53e83ddab2826e999eba3c76f7. - Three fixed seeds:
17,42,73. - One RunPod GPU; each seed is trained and evaluated in the same sealed job.
- LoRA rank 32, alpha 64, dropout 0.05.
- Modules:
q_proj,k_proj,v_proj,out_proj,fc1,fc2. - Six epochs; batch 8; gradient accumulation 2; effective batch 16.
- Learning rate
1e-5; warmup 0.05; weight decay 0.01. - Evaluation/save every 148 optimizer steps (approximately one epoch).
- Best checkpoint within each seed is selected only by Text 36 validation chrF++.
- Full deterministic trainer mode; deterministic float32 inference; TF32 disabled.
- Fixed guarded decoder: one beam, no-repeat 4-gram, repetition penalty 1.10, length penalty 1.0.
- No decoder search in v23.
Seed selection lock
All three seeds must produce 53 nonempty Text 36 predictions with no
tenfold segment loop and mean token length ratio in
[0.5, 2.0]. Among eligible seeds, select maximum Text 36
corpus chrF++; tie-break by mean sentence chrF++, lower maximum
repetition, then lexical seed label. The selection artifact is written
before either baseline or candidate inference on Text 3.
Promotion gates
A research-candidate promotion requires every gate below:
- all three seeds safety-eligible;
- at least two of three seeds improve baseline Text 36 corpus chrF++;
- selected seed improves Text 36 corpus chrF++ by at least 0.5;
- Text 3 corpus chrF++ delta at least +1.0;
- Text 3 paired mean-sentence chrF++ bootstrap 95% interval lower bound above zero;
- more Text 3 row wins than losses;
- positive corpus chrF++ delta on the Text 3 transcription-unflagged slice;
- DB usage corpus chrF++ delta at least -1.0;
- tagged synthetic delta at least -1.0 and untagged synthetic delta at least -1.5;
- elder corpus chrF++ delta at least -1.0;
- dictionary accepted-headword exact count loses no more than three of 297 rows;
- no empty outputs on any evaluated set;
- candidate tenfold-loop count never exceeds baseline;
- Bible direct delta at least -10.0 and no Bible empty output.
The final Bible condition is a catastrophic guard, not positive evidence. A strong Bible score cannot rescue failure on Text 3, replication, elder, usage, dictionary, or degeneration controls.
Statistical interpretation
The paired primary estimand is the v23-minus-v21.2 difference in mean per-row sentence chrF++, with a 50,000-replicate paired row-bootstrap percentile interval. Corpus chrF++ and row win/loss counts are also reported. Rows within a story are not independent speakers, so row-bootstrap intervals do not establish population-level generalization. There is one speaker in each natural split and a deliberate Nyungkul-to-Yalanji transfer step. Promotion means only "best current research candidate"; it does not mean speaker certification, community endorsement, or unrestricted production readiness.
Frozen artifacts
- Dataset manifest SHA-256:
f82f174e4dfbc114f1c7ba6f6c6e14dc87c99a8f4056202a3de11438615e2058 - Training JSONL SHA-256:
6ae80f6525c2bd3d828823cf50aa7bd87dc3ef9d2bcb0b30c90d0c7d01c6c639 - Validation JSONL SHA-256:
34f2c8f28e631a564657e78ba18155ec4b4e0d561d784d1e4d8841cd3c0bd351 - Sealed test JSONL SHA-256:
6d6c1a81ef8be654a8c6d545a70cd5d82a8cc80a63048f0fe658e9513801e6cd - Dataset checksum ledger SHA-256:
e1494cddcf8b200577e6ed808a5a19885beef27df892c4d70b8a1d41fcf61b77 - Runner SHA-256:
f3847e24e54e9f717d9199102c11b409ea1c377c7ecc834cfb0589af9a841643 - Promotion verifier SHA-256:
d89194549ce3fc2d8e8ea4c42e1ea0d39a2e1f4809cd2e69bb388ea99ae6454b
The RunPod input ledger will additionally freeze every uploaded code and model file. Any post-freeze code change requires a new ledger and an explicit amendment before launch.