# Kuku Yalanji v24.3 Exact Training and Evaluation Data

This directory is the exact materialized input package for
`v24.3-joint-lexeme-dose29-s3598-20260715`.

| File | Rows | Purpose |
|---|---:|---|
| `train.eng-gvn.jsonl` | 115,136 | One shuffled pass: 78,996 lexical and 36,140 sentence presentations |
| `validation.eng-gvn.jsonl` | 128 | Mixed optimization monitor; 64 lexical rows overlap training by design |
| `curated-2724.eng-gvn.jsonl` | 2,724 | Exhaustive governed closed-set lexical census |
| `v24-development-suite.eng-gvn.jsonl` | 2,086 | Frozen synthetic, usage, natural-text, elder, and historical-lexicon diagnostics |
| `MANIFEST.json` | - | Source identities, filtering, repetition counts, and leakage audit |
| `run_contract.json` | - | Preregistered causal question, constants, checkpoints, and gates |
| `model_manifest.json` | - | Complete training environment, token inventory, adapter topology, and trainer state |

The training file is a materialized presentation schedule, not 115,136 independent linguistic observations:

- 2,724 governed dictionary records are each shown 29 times;
- 18,070 unique non-Bible sentence rows are each shown twice;
- those sentence rows comprise 16,642 synthetic-academic pairs and 1,428 dictionary-usage pairs;
- no Bible rows are part of the positive v24.3 treatment.

The 2,724-row lexical census overlaps training deliberately and measures closed-set reconstruction. The 43-row
elder-shared diagnostic has been repeatedly inspected and is a regression set, not a blind test. None of these files
is evidence of speaker or community certification.

Verify `SHA256SUMS` after download. Source-specific rights, Indigenous governance requirements, and the upstream
NLLB CC BY-NC 4.0 licence continue to apply; public availability does not erase those obligations.
