Kuku Yalanji v24 Data Requirements and Sentence Commissioning Protocol
Date: 2026-07-15
Program: MobTranslate Kuku Yalanji research
program
Decision status: Data review required; no model
advances to confirmation or deployment
Scope: English to Kuku Yalanji lexical reconstruction,
sentence-data design, and a reusable low-resource MT workflow
Public claim boundary: These experiments do not
establish a reliable free-form translator, speaker certification,
community approval, or authoritative language use.
Executive decision
The next justified action is not another generic synthetic-data run.
The v24 RunPod screen tested increasing doses of explicit grammar-derived lexical pairs. A low dose produced a small gain, while larger doses caused substantial regression. The complete follow-up census then evaluated five frozen candidates over 1,220 grammar-extraction records and 2,724 curated dictionary records. The best treatment, L1, reached only 86/1,220 and 213/2,724 accepted-reference exact reconstructions. It added no exact result at all among the 602 curated prompts whose accepted forms had no identified sentence context.
The evidence supports a narrower conclusion:
- Existing target exposure helps, but exposure alone is insufficient.
- The present
<lexeme>task conflates citation headwords, inflected forms, derivations, compounds, sense-qualified definitions, and ordinary translation. - Larger replay doses encourage related but wrong surface forms, including suffixation, reduplication, compounding, source copying, and selection of a known Kuku Yalanji word for the wrong English prompt.
- The existing 20,047-pair synthetic corpus is broad but highly templated and unevenly distributes lexical and grammatical evidence.
- More data is needed only after source, sense, morphology, variety, rights, and product route have been adjudicated.
The pre-review upper bound is 498 non-name, zero-context prompt records covering 350 unique accepted surface forms that every tested model misses. That is not a requirement for 498 new sentences. A single natural sentence can cover several approved lexical and morphosyntactic cells, while some entries should be lookup-only, excluded, or represented by multiple contrastive sentences. The final sentence count must be produced by the governed review ledger and a researcher-in-the-loop coverage selection, not selected in advance.
Questions tested
The completed study addressed four questions:
- Does direct lexical supervision improve closed-set reconstruction over a step-matched continuation?
- Does increasing lexical dose produce a monotonic improvement?
- Are the original 297 lexical prompts representative of the wider resources?
- Do failures point primarily to insufficient volume, insufficient context, task conflation, tokenization, ambiguity, or morphology?
It did not test whether an arbitrary unseen English-to-Kuku mapping can be inferred. Such a mapping is not productively derivable without prior exposure or a dictionary supplied at inference.
Evidence hierarchy
The report keeps the following evidence classes separate:
| Class | Role | What it can establish |
|---|---|---|
| Governed dictionary record | Deterministic lookup and lexical audit | Attested lexical mapping within its documented sense and variety |
| Speaker/text-attested parallel sentence | Training or held-out natural evaluation, subject to rights | Natural usage within that speaker, text, genre, and context |
| Grammar XIGT extraction | Candidate attested evidence pending source verification | A possible sourced construction or sentence |
| Usage database row | Candidate evidence pending provenance and direction audit | Nothing positive until verified |
| Synthetic pair | Controlled training and regression evidence | Behavior under the documented generation process |
| Model output | Hypothesis only | No linguistic fact without independent review |
No score allows a lower evidence class to overwrite a higher one.
Frozen resources
Historical development suite
The frozen development bundle contains:
| Endpoint | Rows | Interpretation |
|---|---|---|
| Historical closed-set lexicon | 297 | Governed reconstruction diagnostic, not future-query reliability |
| Synthetic development | 1,609 | Same-process retention diagnostic |
| Usage diagnostic | 84 | Unverified usage regression diagnostic |
| Natural development text | 53 | One clustered natural-text diagnostic |
| Elder-shared regression set | 43 | Repeatedly inspected regression set, no longer blind |
Expanded lexical resources
| Resource | Prompt records | Important limitation |
|---|---|---|
| Patz grammar extraction | 1,220 | Heterogeneous LLM extraction; not a curated dictionary and not generation-authorized |
| Curated project dictionary census | 2,724 | Complete resource census, not a random sample of user queries |
| Union of normalized English prompt surfaces | 3,689 | Resources overlap on 255 prompts |
Among the 255 overlapping prompts, 125 share an accepted form, 25 match only after notation removal, 50 have a normalized prefix relation, and 55 remain surface-disjoint. The latter 130 require source and fluent-speaker adjudication before accepted sets can be merged.
RunPod intervention
All continuation arms began from the exact v21.2 lineage and used the same seed, 1,384 optimizer updates, learning-rate trajectory, decoder, and development suite. Actual token exposure was audited rather than inferred from epoch count.
| Label | Intervention |
|---|---|
| B0 | Untouched v21.2 baseline |
| C0 | Step-matched retention-only continuation |
| L1 | Low lexical substitution dose |
| L2 | Medium lexical substitution dose |
| L4 | High lexical substitution dose |
Every continuation arm retained controlled synthetic, usage, and Bible replay. The experiment is therefore a paired lexical-dose mechanism screen, not a proposed no-Bible recipe.
The A40 expanded census used float32 CUDA inference, batch size 32,
greedy decoding, no_repeat_ngram_size=4, repetition penalty
1.10, and the frozen language-token configuration.
Adapter-versus-float32-merge parity was 16/16 for every trained arm.
Frozen-screen result
| Model | Historical exact / 297 | Difference from C0 | Advancement result |
|---|---|---|---|
| B0 | 43 | -2 | Baseline only |
| C0 | 45 | 0 | Control only |
| L1 | 47 | +2 | Fail: required at least +10 |
| L2 | 29 | -16 | Fail |
| L4 | 25 | -20 | Fail |
All sentence-retention noninferiority checks passed, but no treatment
passed the primary lexical threshold. Exact natural-sentence
reconstruction remained 0/53 and exact elder-regression reconstruction
remained 0/43. The selector returned NO_ADVANCE.
Expanded-census result
Accepted-reference exact match gives credit to any reference already authorized for that prompt. It remains a closed-set reconstruction metric, not a sentence-quality metric.
| Model | Grammar extraction exact / 1,220 | Curated census exact / 2,724 |
|---|---|---|
| B0 | 73 (5.98%) | 199 (7.31%) |
| C0 | 75 (6.15%) | 201 (7.38%) |
| L1 | 86 (7.05%) | 213 (7.82%) |
| L2 | 71 (5.82%) | 156 (5.73%) |
| L4 | 70 (5.74%) | 131 (4.81%) |
L1 gained 11 grammar-extraction rows and 12 curated rows over C0. L2 and L4 regressed sharply. L4 also emitted one blank for a dialect-name record. No candidate is suitable for confirmation, mounting, or sentence-generation authorization.
Because the rows are clustered dictionary records rather than a sampled request population, no Wilson interval or percentage in this table should be described as general translation reliability.
Failure anatomy
Curated C0 failures
| Surface category | Rows | Interpretation |
|---|---|---|
| Accepted exact | 201 | Correct closed-set reconstruction |
| Accepted form plus extra material | 50 | Accepted form is present, but output is not the required citation response |
| Known Kuku target for another prompt | 809 | Target-like output with wrong source alignment |
| Near orthographic | 41 | Close surface form requiring source/linguistic review |
| Notation or hyphen mismatch | 5 | Representation issue, not automatically a lexical error |
| Source copy | 52 | English or prompt material copied into output |
| Other failure | 1,566 | No accepted surface relation found |
L1 reduces other failures but increases
wrong-known-target outputs from 809 to 875. Its 20 curated gains and
eight losses are not broad unseen-lexeme acquisition. Representative
gains remove unwanted material from already contextualized forms:
water: bana-ji -> bana
foot: jina-ji -> jina
know: binalku -> binal
hot coals: ngunjil wumbul -> ngunjil
young girl: maral buban -> maral
Representative losses include:
above: wangkar-wangkar -> wangkar-a
quiet: janka -> janka-janka
underneath: bada-bada -> bada-ba
darkness: wujurr -> ngiki-ngki
hunt: balbi -> bayka
These outputs show citation-form confusion, over-expansion, lexical competition, and semantic substitution. They do not support simply increasing replay dose.
Context exposure
The evidence-linked queue joins every curated prompt to synthetic, usage, and source-backed XIGT context.
| Cohort | Prompt records | L1 accepted exact | Every model fails exact |
|---|---|---|---|
| Some identified context | 2,122 | 198 (9.33%) | 1,906 |
| No identified context | 602 | 15 (2.49%) | 587 |
| Linked to synthetic context | 1,898 | 197 (10.38%) | 1,684 |
| Linked to source-backed XIGT context | 310 | 77 (24.84%) | 226 |
| Linked to unverified usage context | 641 | 35 (5.46%) | 598 |
The cohorts overlap. Prompt counts are larger than unique-headword counts because a headword can support several English prompts and a prompt can accept several forms.
Most importantly, L1 adds zero exact results over C0 among the 602 zero-context prompts. Its gains occur only where some context was already identified. Direct replay of the current grammar extraction therefore did not solve the missing-context problem.
Exposure frequency
For C0, curated reconstruction rises from 24/947 (2.53%) when the selected target has zero explicit synthetic occurrences to 109/620 (17.58%) at 20 or more occurrences. Exposure matters, but even the highest frequency band fails more than four times in five. Frequency is neither a substitute for sense alignment nor evidence of grammatical use.
English prompt shape
| English prompt length | C0 exact | Rate |
|---|---|---|
| 0-1 word | 185/1,563 | 11.84% |
| 2 words | 13/794 | 1.64% |
| 3-5 words | 3/341 | 0.88% |
| 6+ words | 0/26 | 0.00% |
Longer dictionary definitions and sense descriptions are not behaving like isolated-word prompts. They need explicit sense/POS conditioning and separate evaluation.
Part of speech
| Part of speech | C0 exact | Rate |
|---|---|---|
| Noun | 144/1,616 | 8.91% |
| Adjective | 26/382 | 6.81% |
| Intransitive verb | 16/277 | 5.78% |
| Transitive verb | 7/376 | 1.86% |
The particularly weak transitive-verb result means the missing evidence is not merely a word list. Verb supervision must represent participants, valency, case, TAM, polarity, and clause environment.
Tokenization
C0 exact rates are 7.69% for targets represented by 0-1 counted subwords, 8.76% for two, 6.46% for 3-4, and 5.47% for 5-8. Fragmentation contributes, especially at the tail, but it is much smaller than the source-shape and exposure effects.
A collision-free tokenizer-extension prototype reduced target fertility by 12.44% on its synthetic-skewed fit corpus but only 3.37% on the curated census and affected just 9.47% of curated records. It is worth an isolated future arm, not a sufficient diagnosis.
Does the program need more synthetic sentences?
Not more undifferentiated sentences from the existing process.
The current corpus already links synthetic context to 1,898 curated prompts, yet every tested model fails 1,684 of them. This can reflect inadequate or wrong-sense contexts, sparse forms, task mismatch, templatic repetition, frozen target-language embeddings, or model capacity. Simply increasing the same generator's volume would confound these causes and repeat the failed dose logic.
New bilingual sentence evidence is justified only for a reviewed coverage cell. Preferred evidence order is:
- Reuse and verify an existing speaker/text-attested sentence.
- Verify a source-backed grammar XIGT sentence.
- Audit and repair an existing usage row.
- Commission a new sentence from a fluent speaker or translator under documented governance.
- Use an inspected English-side generated candidate only as an elicitation prompt; the Kuku Yalanji target must not be represented as speaker-attested unless a fluent speaker supplies or approves it.
Required work before another training run
Gate 1: sense-level dictionary adjudication
Review all 2,724 prompt records and retain stable identifiers for:
English form
part of speech
sense identifier
accepted Kuku Yalanji citation form(s)
inflection/derivation status
variety and register
source and page/entry citation
rights and governance state
product route: lookup, model research, restricted, or excluded
Resolve the 476 ambiguous prompts and the 130 cross-resource conflicts before building accepted-reference sets. The 57 identity-like mappings require explicit source confirmation rather than automatic approval.
Gate 2: salvage existing attested candidates
Before commissioning new material:
- Validate the 295 source-backed XIGT clause candidates against the grammar pages and preserve their surrounding discourse. These map to 310 curated prompt records and 182 headwords because records overlap.
- Audit all 449 usage rows (446 unique pairs). Their current provenance is incomplete, and all are marked unverified; at least one historical replay row is directionally reversed.
- Review existing synthetic contexts against the cited grammar rule, intended sense, target form, and clause structure instead of trusting a surface headword occurrence.
- Mark every retained row as train-only, development, final test, or excluded before model fitting.
Gate 3: remove lookup-only and restricted cells
The curated inventory contains 572 zero-context headwords. Of these,
133 are in *-name semantic domains such as place, personal,
or tribal-group names. These normally belong in deterministic dictionary
or metadata lookup and should not consume sentence-model capacity unless
a governed research question requires their contextual use.
Ceremonial, spiritual, culturally sensitive, or otherwise restricted entries require the relevant authority and use conditions. Public availability alone is not automatic authorization for model training, hosted-provider transfer, or model-weight distribution.
Gate 4: construct the sentence-commissioning inventory
After Gates 1-3, begin from the current maximum candidate pool:
| Pre-review pool | Prompt records | Unique accepted surface forms |
|---|---|---|
| Zero context and every model fails | 587 | 446 |
| Name-domain subset | 89 | 97 |
| Non-name subset | 498 | 350 |
The 498/350 non-name figures are upper bounds, not a final quota. Remove wrong senses, duplicates, lookup-only items, restricted material, and cells already covered by validated XIGT or usage evidence.
Gate 5: commission coverage, not one sentence per word
Each approved sentence candidate should cover one or more explicit cells:
LEX-CONTEXT: lexical sense in natural context
- one governed sense and POS;
- a natural source sentence where that sense is unambiguous;
- the appropriate target citation stem and any required morphology;
- no forced insertion of an isolated headword into an unnatural sentence.
VALENCY: argument structure
- transitive versus intransitive use;
- agent and patient reversal;
- overt and contextually omitted arguments;
- relevant case frames;
- passive, antipassive, or other documented alternations where source-backed and appropriate.
MORPH: productive morphology
- lemma observed in training, with held-out attested inflected or derived forms;
- case, TAM, polarity, number, derivation, clitic/particle, and reduplication cells supported by the grammar;
- minimal or contrastive families reviewed as natural Kuku Yalanji, not mechanical paradigm expansion.
CLAUSE: sentence type
- declarative, interrogative, imperative, negative, subordinate, and linked-clause environments;
- information-structure and word-order contrasts where documented;
- length and participant structure representative of intended use.
DISCOURSE: independent natural text
- contiguous discourse rather than isolated clauses only;
- speaker, text, genre, occasion, variety, and context identifiers;
- independent clusters reserved for development and final evaluation;
- coreference, topic continuity, ellipsis, and discourse particles retained.
GLOSSARY-UPTAKE: inference-time new lexical knowledge
Hold a governed mapping out of model training, provide it at inference, and test whether the sentence model uses it appropriately:
<translate> The woman returned.
<glossary> woman = jalbu
This is the meaningful unseen-lexeme task. Asking a model to guess an undisclosed arbitrary mapping is not productive lexical generalization.
Coverage-driven stopping rule
The final sentence count is calculated after adjudication using a reviewed coverage matrix.
- Define the approved universe
Uof lexical-sense, morphology, valency, clause, variety, and discourse cells. - Mark cells already covered by governed training evidence and cells reserved only for evaluation.
- Generate or retrieve natural English-side candidates that may cover uncovered cells.
- Have a researcher inspect and edit every source candidate. Reject short content-word piles and unnatural coverage "honeypots."
- Obtain fluent-speaker Kuku Yalanji translation or approval under the recorded governance conditions.
- Score each candidate by newly covered cells, review cost, naturalness, cluster independence, and duplication.
- Select iteratively until no eligible high-priority cell remains uncovered or the governing reviewers defer it.
- Freeze family/speaker/text-disjoint development and final sets before training.
This adopts the researcher-in-the-loop set-cover principle demonstrated by SMOLSENT without copying its English inventory or its 863-sentence count. Kuku Yalanji's final number must emerge from its own governed inventory.
Required row schema
Every retained bilingual sentence should carry at least:
{
"id": "immutable-row-id",
"source_text": "English source",
"target_text": "Kuku Yalanji target",
"origin": "speaker_authored_or_grammar_attested_or_reviewed_elicitation",
"speaker_cluster_id": "stable-or-protected-id",
"text_cluster_id": "stable-id",
"source_citation": "work-page-line-or-recording-span",
"lexeme_sense_ids": ["stable-sense-id"],
"morphosyntactic_cells": ["documented-cell-id"],
"variety": "documented-variety-or-unknown",
"register": "documented-register-or-unknown",
"rights": {
"training": false,
"redistribution": false,
"hosted_provider_transfer": false,
"model_weight_distribution": false
},
"review": {
"fluent_review_status": "pending",
"reviewer_authority": "recorded-separately",
"review_date": null
},
"split_lock": "train_or_dev_or_final_or_excluded"
}Unknown values must remain explicit. Missing provenance must not silently become approval.
Next model experiment after the data gates
The next model study should be narrower than the completed screen.
- Keep deterministic dictionary lookup as the production route for governed lexical queries.
- Use separate model-visible tasks:
<lexeme> woman
<translate> The woman returned.
<translate> The woman returned. <glossary> woman = jalbu
<morphology> [documented structured condition]
- Use a clean no-Bible positive objective. A frozen Bible set may remain a catastrophic-forgetting diagnostic but cannot select a checkpoint or offset failure elsewhere.
- Audit and train the custom
gvn_Latnand task-control token rows. In the current lineage, the copiedgvn_Latnembedding remained bit-identical to its initialization because embeddings were frozen. - Test the collision-free tokenizer extension as a separate factor after the trainable-language-token arm, not simultaneously.
- Use exact
max_steps, paired seeds, fixed token budgets, fixed learning-rate integrals, and per-task token accounting. - Keep an untouched baseline and a token-accounted continuation control.
- Select a recipe on aggregate development evidence, then run multiple confirmatory seeds only if the screen passes.
- Keep lexical reconstruction, glossary uptake, morphology, sentence adequacy, and decoder degeneration as separate endpoints.
- Do not let a dictionary score authorize sentence deployment.
NLLB/Formosan tutorial disposition
The project read and audited the complete FormosanBank NLLB adaptation tutorial rather than copying its commands mechanically.
Adopted ideas:
- directional corpora and explicit source/target contracts;
- leakage-controlled source, speaker, and text splits;
- one-token language and task controls;
- training a SentencePiece candidate only on training data;
- initializing new embeddings from decomposed inherited pieces;
- explicit decoder start/EOS behavior;
- preserving a reproducible tokenizer/model bundle.
Material correction:
Literal merging of new SentencePiece pieces into NLLB can overlap the
inherited language-token ID range. The project reproduced this failure:
candidate lexical IDs decoded as unrelated language tags such as
kon_Latn and acm_Arab. The published tutorial
tokenizer shows the same structural range risk. MobTranslate therefore
built a collision-free extension that preserves every inherited token ID
and appends new pieces and controls after the reserved range. That
prototype passes its invariants but has not been used to select or train
a model.
Reusable protocol for another Aboriginal language
This workflow is intended to transfer, while the linguistic decisions and governance do not.
- Governance and rights: record authority, scope, restrictions, benefit, withdrawal, hosted transfer, and weight-release conditions.
- Sense ledger: assign stable lexical sense/POS/variety/citation-form identifiers; never collapse by English spelling alone.
- Evidence inventory: separate dictionary, natural text, grammar examples, usage rows, synthetic rows, and model hypotheses.
- Split lock: freeze source-, speaker-, text-, and lexeme-family-disjoint development/final material before fitting.
- Baseline routes: benchmark deterministic lookup separately from model reconstruction and sentence translation.
- Full-resource census: run all candidate models over complete lexical resources, but label it a census rather than a population sample.
- Qualitative error taxonomy: measure exact, accepted-plus-extra, orthographic, notation, wrong-known-target, source-copy, blank, and uncategorized failures.
- Context join: connect every lexical sense to existing natural, grammar, usage, and synthetic contexts.
- Salvage first: verify attested candidates before commissioning new material.
- Coverage selection: build the language-specific lexical, morphosyntactic, variety, and discourse universe; use researcher-in-the-loop set cover.
- Speaker review: obtain or approve target sentences through the relevant fluent speakers and authorities; record uncertainty and disagreement.
- Tokenizer/language-token audit: verify ID preservation, fertility, trainability, tied embeddings, adapter persistence, and merged parity.
- Controlled screen: vary one causal factor at a time under matched updates, tokens, and learning-rate exposure.
- Confirmatory evaluation: use independent clusters, multiple seeds, paired analysis, and a once-opened final test.
- Split release gates: lookup, lexical reconstruction, and sentence generation receive independent decisions.
- Artifact policy: retain compact adapters, manifests, predictions, code, data order, and hashes; delete redundant merged failures and optimizer checkpoints after verified preservation.
Reproduction commands
The expanded census was run with:
WORK_ROOT=/workspace/v24-lexicon-screen \
training/translation/run_v24_expanded_lexicon_census.shThe evidence queue was built with:
PROG=/mnt/donto-data/donto-resources/research/translation-training/kuku-yalanji-runpod-2026-06-30
V24="$PROG/runpod/v24.0-lexicon-grounded-screen-20260715T044520Z"
EXPANDED="$V24/analysis/expanded-lexicon-census-20260715"
python3 training/translation/build_sentence_evidence_review_queue.py \
--curated-census "$PROG/prepared/kuku-yalanji-curated-dictionary-census-20260715.eng-gvn.jsonl" \
--context-inventory "$PROG/analysis/lexeme-context-inventory-20260715/inventory.jsonl" \
--curated-row-analysis "$EXPANDED/analysis/curated-row-analysis.jsonl" \
--overlap-rows "$V24/analysis/lexicon-overlap-audit-20260715/overlaps.jsonl" \
--expected-candidate B0 --expected-candidate C0 --expected-candidate L1 \
--expected-candidate L2 --expected-candidate L4 \
--output "$PROG/analysis/sentence-evidence-review-queue-20260715/queue.jsonl" \
--summary-output "$PROG/analysis/sentence-evidence-review-queue-20260715/summary.json"Integrity and retained artifacts
| Artifact | SHA-256 |
|---|---|
| Frozen selector | d4a3623c60bcbce804e132f4ff7c422b3298ea72119a4dcdbbb6d6e896693c0b |
| Expanded final checksum ledger | 60b127c25d18d56fb0e8b415502981d17998ee7bc43ca3a912e7a8f64569a95b |
| Expanded raw summary | 74d716346a77a902fb006ee3cffafdb4047517ecdf825ed0bdb84c190e43beb0 |
| Expanded curated summary | b7742a914398733a8d8ce6142fb7acc380839a34379e441574dbbb815260da63 |
| Context inventory | ea1ed3c7d8bcbbc484ebc257d3c50d72134d5d4b9fec1e7a7d610e8fea898272 |
| Review queue | fca2ba770ebf67358573e8854480d76356b43f06c0ccdfc85d14c68db3b39923 |
| Queue summary | b8c18ff0ac10de92a47fb4edad44aae2fc156eca927ebd519aa45f9c833fe79c |
| Compact v24 artifact ledger | 7f615cadb78363f2c0f7c3437cce6c7e5e7d0933b0da159b0884e15332a6214d |
The compact archive retains exact arm datasets, adapters, tokenizers,
manifests, parity checks, predictions, analyses, logs, and 244-file
checksums. Failed merged models, optimizer states, and redundant
checkpoints were not retained. The A40 pod was stopped and deleted after
local hash verification; runpodctl pod list returned an
empty list.
Decision table
| Decision | Result |
|---|---|
| Train another model immediately | No |
| Generate another generic 20,000 synthetic pairs | No |
| Treat L1 as a successful model | No |
| Use L1 as evidence about task/data design | Yes |
| Expand the benchmark from 297 prompts | Completed: 1,220 + 2,724 resource censuses |
| Qualitatively classify failures | Completed and joined to context evidence |
| Commission 498 new sentences automatically | No |
| Review an upper-bound pool of 498 non-name prompts / 350 forms | Yes |
| Validate existing XIGT and usage candidates first | Yes |
| Use deterministic dictionary lookup | Yes |
| Mount a custom Kuku Yalanji sentence model | No |
References
- Elisabeth Patz, A Grammar of the Kuku Yalanji Language of North Queensland: https://openresearch-repository.anu.edu.au/items/69491755-3b9d-4323-a3ef-2e6ac075042f
- FormosanBank, Training NLLB-200 for a New Language: https://huggingface.co/blog/FormosanBank/nllb-200-mt
- Meta,
facebook/nllb-200-distilled-1.3Bmodel card: https://huggingface.co/facebook/nllb-200-distilled-1.3B - GATITOS bilingual lexicon augmentation study: https://aclanthology.org/2023.emnlp-main.26/
- SMOL low-resource parallel-data selection study: https://aclanthology.org/2025.wmt-1.85/
- SacreBLEU reproducible MT evaluation: https://aclanthology.org/W18-6319/
- AIATSIS Code of Ethics: https://aiatsis.gov.au/sites/default/files/2020-10/aiatsis-code-ethics.pdf
- CARE Principles for Indigenous Data Governance: https://www.gida-global.org/careprinciples
Final conclusion
The expanded benchmark answered the immediate question. More exposure is associated with better lexical reconstruction, but the current model does not convert generic volume into reliable source-target alignment, morphology, or sentence competence. Increasing lexical replay beyond L1 is actively harmful, and L1 does not improve the zero-context cohort.
The next useful corpus is therefore not "another N synthetic sentences." It is the smallest governed set of natural, fluent-reviewed bilingual sentences that remains after sense adjudication, attested-evidence salvage, lookup routing, and coverage selection. The review ledger now identifies the maximum candidate pool and the exact decisions needed to turn it into that set. Training should resume only when those decisions produce a checksum-frozen, split-locked data manifest.