mobtranslate.com / docs › Kuku Yalanji source: kuku-v23-preregistration.md

Kuku Yalanji v23 Attested-Narrative Adaptation Preregistration

Status: frozen before RunPod training or v23 inference
Date: 2026-07-14 UTC
Experiment: v23.0-attested-narrative-adaptation

Research question

Does a low-learning-rate second-stage LoRA over the retained v21.2 model improve English-to-Kuku Yalanji translation on source-attested, speaker-disjoint natural narrative, without materially damaging everyday usage, dictionary coverage, the elder diagnostic, or the existing synthetic competence?

The treatment is not another exposure-length point on the v21.2 synthetic/Bible trajectory. It adds a manually reconciled native-speaker narrative corpus and uses restrained retention replay. The Bible has zero training rows, zero seed-selection weight, and zero promotion-objective weight. One Bible set is run only as a loose catastrophic-forgetting alarm.

Evidential hierarchy

  1. Primary sealed test: Patz Text 3, Ivy Walker, Yalanji, 56 clauses. This is opened only after seed selection has been written to seed_selection.locked.json.
  2. Development and replication: Patz Text 36, Bobby Roberts, Nyungkul, 53 clauses. This alone selects checkpoints and one of three training seeds.
  3. Culturally important nonblind diagnostic: 43 rights-cleared elder sentence pairs. These rows have been inspected in earlier work and are not represented as a blind test.
  4. Behavioral controls: 84 held-out database usage examples, the 297-row multi-reference dictionary probe, and the frozen tagged/untagged synthetic tests.
  5. Bible: one 325-row direct-translation set, used only to detect catastrophic forgetting at a deliberately broad minus-10 chrF++ margin. It cannot select a seed or compensate for a natural-test failure.

No metric is pooled across these strata. In particular, Bible volume cannot dominate model judgment.

Frozen corpus

The deterministic builder is training/translation/build_v23_attested_adaptation.py.

Split or component Rows Role
Texts 51 and 12, Charlie Tayley, Nyungkul 156 unique / 624 after fixed x4 replay New attested train signal
DB usage, direct mode 356 Everyday/project-database train signal
DB usage, original glossary mode 356 v21 task-mode retention
Synthetic retention sample 1,024 Existing competence retention
Training total 2,360 Six epochs per seed
Text 36, Bobby Roberts, Nyungkul 53 Validation and seed selection
Text 3, Ivy Walker, Yalanji 56 Sealed natural test
Bible 0 train rows No training role

Nine Text 51/12 clauses carrying explicit transcription/source uncertainty warnings are excluded from training. A tenth candidate clause, 51.87 (Yuwu, "Yes"), is quarantined because its normalized target duplicates held-out 3.52. Nine DB rows and six synthetic rows are quarantined by the frozen exact and token-overlap rules. DB word IDs are disjoint from the 84-row heldout.

Why the 795-row XIGT export is not training data in v23

All 795 records are retained in audit/xigt_exclusion_audit.json, but none is approved for this run. The export mixes words, paradigms, phrases, clauses, unresolved target/gloss interleaving, and known systematic OCR substitutions (l/J/1). A mechanical screen labels 557 rows as superficially clean, but that does not establish correct Kuku orthography or clause-level verification. Calling all 795 records "attested sentence pairs" would overstate the evidence and risk teaching OCR errors. A later version may use them only after independent normalization and review.

Leakage controls

Before training, candidate training rows are compared with the natural validation/test and inherited heldouts using:

Every exclusion and reason is additive in audit/quarantine.json. No held-out or test file is passed to train_nllb_lora.py.

Model and training treatment

Seed selection lock

All three seeds must produce 53 nonempty Text 36 predictions with no tenfold segment loop and mean token length ratio in [0.5, 2.0]. Among eligible seeds, select maximum Text 36 corpus chrF++; tie-break by mean sentence chrF++, lower maximum repetition, then lexical seed label. The selection artifact is written before either baseline or candidate inference on Text 3.

Promotion gates

A research-candidate promotion requires every gate below:

The final Bible condition is a catastrophic guard, not positive evidence. A strong Bible score cannot rescue failure on Text 3, replication, elder, usage, dictionary, or degeneration controls.

Statistical interpretation

The paired primary estimand is the v23-minus-v21.2 difference in mean per-row sentence chrF++, with a 50,000-replicate paired row-bootstrap percentile interval. Corpus chrF++ and row win/loss counts are also reported. Rows within a story are not independent speakers, so row-bootstrap intervals do not establish population-level generalization. There is one speaker in each natural split and a deliberate Nyungkul-to-Yalanji transfer step. Promotion means only "best current research candidate"; it does not mean speaker certification, community endorsement, or unrestricted production readiness.

Frozen artifacts

The RunPod input ledger will additionally freeze every uploaded code and model file. Any post-freeze code change requires a new ledger and an explicit amendment before launch.