mobtranslate.com / docs › Kuku Yalanji source: kuku-v24-data-requirements.md

Kuku Yalanji v24 Data Requirements and Sentence Commissioning Protocol

Date: 2026-07-15
Program: MobTranslate Kuku Yalanji research program
Decision status: Data review required; no model advances to confirmation or deployment
Scope: English to Kuku Yalanji lexical reconstruction, sentence-data design, and a reusable low-resource MT workflow
Public claim boundary: These experiments do not establish a reliable free-form translator, speaker certification, community approval, or authoritative language use.

Executive decision

The next justified action is not another generic synthetic-data run.

The v24 RunPod screen tested increasing doses of explicit grammar-derived lexical pairs. A low dose produced a small gain, while larger doses caused substantial regression. The complete follow-up census then evaluated five frozen candidates over 1,220 grammar-extraction records and 2,724 curated dictionary records. The best treatment, L1, reached only 86/1,220 and 213/2,724 accepted-reference exact reconstructions. It added no exact result at all among the 602 curated prompts whose accepted forms had no identified sentence context.

The evidence supports a narrower conclusion:

  1. Existing target exposure helps, but exposure alone is insufficient.
  2. The present <lexeme> task conflates citation headwords, inflected forms, derivations, compounds, sense-qualified definitions, and ordinary translation.
  3. Larger replay doses encourage related but wrong surface forms, including suffixation, reduplication, compounding, source copying, and selection of a known Kuku Yalanji word for the wrong English prompt.
  4. The existing 20,047-pair synthetic corpus is broad but highly templated and unevenly distributes lexical and grammatical evidence.
  5. More data is needed only after source, sense, morphology, variety, rights, and product route have been adjudicated.

The pre-review upper bound is 498 non-name, zero-context prompt records covering 350 unique accepted surface forms that every tested model misses. That is not a requirement for 498 new sentences. A single natural sentence can cover several approved lexical and morphosyntactic cells, while some entries should be lookup-only, excluded, or represented by multiple contrastive sentences. The final sentence count must be produced by the governed review ledger and a researcher-in-the-loop coverage selection, not selected in advance.

Questions tested

The completed study addressed four questions:

  1. Does direct lexical supervision improve closed-set reconstruction over a step-matched continuation?
  2. Does increasing lexical dose produce a monotonic improvement?
  3. Are the original 297 lexical prompts representative of the wider resources?
  4. Do failures point primarily to insufficient volume, insufficient context, task conflation, tokenization, ambiguity, or morphology?

It did not test whether an arbitrary unseen English-to-Kuku mapping can be inferred. Such a mapping is not productively derivable without prior exposure or a dictionary supplied at inference.

Evidence hierarchy

The report keeps the following evidence classes separate:

Class Role What it can establish
Governed dictionary record Deterministic lookup and lexical audit Attested lexical mapping within its documented sense and variety
Speaker/text-attested parallel sentence Training or held-out natural evaluation, subject to rights Natural usage within that speaker, text, genre, and context
Grammar XIGT extraction Candidate attested evidence pending source verification A possible sourced construction or sentence
Usage database row Candidate evidence pending provenance and direction audit Nothing positive until verified
Synthetic pair Controlled training and regression evidence Behavior under the documented generation process
Model output Hypothesis only No linguistic fact without independent review

No score allows a lower evidence class to overwrite a higher one.

Frozen resources

Historical development suite

The frozen development bundle contains:

Endpoint Rows Interpretation
Historical closed-set lexicon 297 Governed reconstruction diagnostic, not future-query reliability
Synthetic development 1,609 Same-process retention diagnostic
Usage diagnostic 84 Unverified usage regression diagnostic
Natural development text 53 One clustered natural-text diagnostic
Elder-shared regression set 43 Repeatedly inspected regression set, no longer blind

Expanded lexical resources

Resource Prompt records Important limitation
Patz grammar extraction 1,220 Heterogeneous LLM extraction; not a curated dictionary and not generation-authorized
Curated project dictionary census 2,724 Complete resource census, not a random sample of user queries
Union of normalized English prompt surfaces 3,689 Resources overlap on 255 prompts

Among the 255 overlapping prompts, 125 share an accepted form, 25 match only after notation removal, 50 have a normalized prefix relation, and 55 remain surface-disjoint. The latter 130 require source and fluent-speaker adjudication before accepted sets can be merged.

RunPod intervention

All continuation arms began from the exact v21.2 lineage and used the same seed, 1,384 optimizer updates, learning-rate trajectory, decoder, and development suite. Actual token exposure was audited rather than inferred from epoch count.

Label Intervention
B0 Untouched v21.2 baseline
C0 Step-matched retention-only continuation
L1 Low lexical substitution dose
L2 Medium lexical substitution dose
L4 High lexical substitution dose

Every continuation arm retained controlled synthetic, usage, and Bible replay. The experiment is therefore a paired lexical-dose mechanism screen, not a proposed no-Bible recipe.

The A40 expanded census used float32 CUDA inference, batch size 32, greedy decoding, no_repeat_ngram_size=4, repetition penalty 1.10, and the frozen language-token configuration. Adapter-versus-float32-merge parity was 16/16 for every trained arm.

Frozen-screen result

Model Historical exact / 297 Difference from C0 Advancement result
B0 43 -2 Baseline only
C0 45 0 Control only
L1 47 +2 Fail: required at least +10
L2 29 -16 Fail
L4 25 -20 Fail

All sentence-retention noninferiority checks passed, but no treatment passed the primary lexical threshold. Exact natural-sentence reconstruction remained 0/53 and exact elder-regression reconstruction remained 0/43. The selector returned NO_ADVANCE.

Expanded-census result

Accepted-reference exact match gives credit to any reference already authorized for that prompt. It remains a closed-set reconstruction metric, not a sentence-quality metric.

Model Grammar extraction exact / 1,220 Curated census exact / 2,724
B0 73 (5.98%) 199 (7.31%)
C0 75 (6.15%) 201 (7.38%)
L1 86 (7.05%) 213 (7.82%)
L2 71 (5.82%) 156 (5.73%)
L4 70 (5.74%) 131 (4.81%)

L1 gained 11 grammar-extraction rows and 12 curated rows over C0. L2 and L4 regressed sharply. L4 also emitted one blank for a dialect-name record. No candidate is suitable for confirmation, mounting, or sentence-generation authorization.

Because the rows are clustered dictionary records rather than a sampled request population, no Wilson interval or percentage in this table should be described as general translation reliability.

Failure anatomy

Curated C0 failures

Surface category Rows Interpretation
Accepted exact 201 Correct closed-set reconstruction
Accepted form plus extra material 50 Accepted form is present, but output is not the required citation response
Known Kuku target for another prompt 809 Target-like output with wrong source alignment
Near orthographic 41 Close surface form requiring source/linguistic review
Notation or hyphen mismatch 5 Representation issue, not automatically a lexical error
Source copy 52 English or prompt material copied into output
Other failure 1,566 No accepted surface relation found

L1 reduces other failures but increases wrong-known-target outputs from 809 to 875. Its 20 curated gains and eight losses are not broad unseen-lexeme acquisition. Representative gains remove unwanted material from already contextualized forms:

water:      bana-ji       -> bana
foot:       jina-ji       -> jina
know:       binalku       -> binal
hot coals:  ngunjil wumbul -> ngunjil
young girl: maral buban   -> maral

Representative losses include:

above:      wangkar-wangkar -> wangkar-a
quiet:      janka           -> janka-janka
underneath: bada-bada       -> bada-ba
darkness:   wujurr          -> ngiki-ngki
hunt:       balbi           -> bayka

These outputs show citation-form confusion, over-expansion, lexical competition, and semantic substitution. They do not support simply increasing replay dose.

Context exposure

The evidence-linked queue joins every curated prompt to synthetic, usage, and source-backed XIGT context.

Cohort Prompt records L1 accepted exact Every model fails exact
Some identified context 2,122 198 (9.33%) 1,906
No identified context 602 15 (2.49%) 587
Linked to synthetic context 1,898 197 (10.38%) 1,684
Linked to source-backed XIGT context 310 77 (24.84%) 226
Linked to unverified usage context 641 35 (5.46%) 598

The cohorts overlap. Prompt counts are larger than unique-headword counts because a headword can support several English prompts and a prompt can accept several forms.

Most importantly, L1 adds zero exact results over C0 among the 602 zero-context prompts. Its gains occur only where some context was already identified. Direct replay of the current grammar extraction therefore did not solve the missing-context problem.

Exposure frequency

For C0, curated reconstruction rises from 24/947 (2.53%) when the selected target has zero explicit synthetic occurrences to 109/620 (17.58%) at 20 or more occurrences. Exposure matters, but even the highest frequency band fails more than four times in five. Frequency is neither a substitute for sense alignment nor evidence of grammatical use.

English prompt shape

English prompt length C0 exact Rate
0-1 word 185/1,563 11.84%
2 words 13/794 1.64%
3-5 words 3/341 0.88%
6+ words 0/26 0.00%

Longer dictionary definitions and sense descriptions are not behaving like isolated-word prompts. They need explicit sense/POS conditioning and separate evaluation.

Part of speech

Part of speech C0 exact Rate
Noun 144/1,616 8.91%
Adjective 26/382 6.81%
Intransitive verb 16/277 5.78%
Transitive verb 7/376 1.86%

The particularly weak transitive-verb result means the missing evidence is not merely a word list. Verb supervision must represent participants, valency, case, TAM, polarity, and clause environment.

Tokenization

C0 exact rates are 7.69% for targets represented by 0-1 counted subwords, 8.76% for two, 6.46% for 3-4, and 5.47% for 5-8. Fragmentation contributes, especially at the tail, but it is much smaller than the source-shape and exposure effects.

A collision-free tokenizer-extension prototype reduced target fertility by 12.44% on its synthetic-skewed fit corpus but only 3.37% on the curated census and affected just 9.47% of curated records. It is worth an isolated future arm, not a sufficient diagnosis.

Does the program need more synthetic sentences?

Not more undifferentiated sentences from the existing process.

The current corpus already links synthetic context to 1,898 curated prompts, yet every tested model fails 1,684 of them. This can reflect inadequate or wrong-sense contexts, sparse forms, task mismatch, templatic repetition, frozen target-language embeddings, or model capacity. Simply increasing the same generator's volume would confound these causes and repeat the failed dose logic.

New bilingual sentence evidence is justified only for a reviewed coverage cell. Preferred evidence order is:

  1. Reuse and verify an existing speaker/text-attested sentence.
  2. Verify a source-backed grammar XIGT sentence.
  3. Audit and repair an existing usage row.
  4. Commission a new sentence from a fluent speaker or translator under documented governance.
  5. Use an inspected English-side generated candidate only as an elicitation prompt; the Kuku Yalanji target must not be represented as speaker-attested unless a fluent speaker supplies or approves it.

Required work before another training run

Gate 1: sense-level dictionary adjudication

Review all 2,724 prompt records and retain stable identifiers for:

English form
part of speech
sense identifier
accepted Kuku Yalanji citation form(s)
inflection/derivation status
variety and register
source and page/entry citation
rights and governance state
product route: lookup, model research, restricted, or excluded

Resolve the 476 ambiguous prompts and the 130 cross-resource conflicts before building accepted-reference sets. The 57 identity-like mappings require explicit source confirmation rather than automatic approval.

Gate 2: salvage existing attested candidates

Before commissioning new material:

  1. Validate the 295 source-backed XIGT clause candidates against the grammar pages and preserve their surrounding discourse. These map to 310 curated prompt records and 182 headwords because records overlap.
  2. Audit all 449 usage rows (446 unique pairs). Their current provenance is incomplete, and all are marked unverified; at least one historical replay row is directionally reversed.
  3. Review existing synthetic contexts against the cited grammar rule, intended sense, target form, and clause structure instead of trusting a surface headword occurrence.
  4. Mark every retained row as train-only, development, final test, or excluded before model fitting.

Gate 3: remove lookup-only and restricted cells

The curated inventory contains 572 zero-context headwords. Of these, 133 are in *-name semantic domains such as place, personal, or tribal-group names. These normally belong in deterministic dictionary or metadata lookup and should not consume sentence-model capacity unless a governed research question requires their contextual use.

Ceremonial, spiritual, culturally sensitive, or otherwise restricted entries require the relevant authority and use conditions. Public availability alone is not automatic authorization for model training, hosted-provider transfer, or model-weight distribution.

Gate 4: construct the sentence-commissioning inventory

After Gates 1-3, begin from the current maximum candidate pool:

Pre-review pool Prompt records Unique accepted surface forms
Zero context and every model fails 587 446
Name-domain subset 89 97
Non-name subset 498 350

The 498/350 non-name figures are upper bounds, not a final quota. Remove wrong senses, duplicates, lookup-only items, restricted material, and cells already covered by validated XIGT or usage evidence.

Gate 5: commission coverage, not one sentence per word

Each approved sentence candidate should cover one or more explicit cells:

LEX-CONTEXT: lexical sense in natural context

VALENCY: argument structure

MORPH: productive morphology

CLAUSE: sentence type

DISCOURSE: independent natural text

GLOSSARY-UPTAKE: inference-time new lexical knowledge

Hold a governed mapping out of model training, provide it at inference, and test whether the sentence model uses it appropriately:

<translate> The woman returned.
<glossary> woman = jalbu

This is the meaningful unseen-lexeme task. Asking a model to guess an undisclosed arbitrary mapping is not productive lexical generalization.

Coverage-driven stopping rule

The final sentence count is calculated after adjudication using a reviewed coverage matrix.

  1. Define the approved universe U of lexical-sense, morphology, valency, clause, variety, and discourse cells.
  2. Mark cells already covered by governed training evidence and cells reserved only for evaluation.
  3. Generate or retrieve natural English-side candidates that may cover uncovered cells.
  4. Have a researcher inspect and edit every source candidate. Reject short content-word piles and unnatural coverage "honeypots."
  5. Obtain fluent-speaker Kuku Yalanji translation or approval under the recorded governance conditions.
  6. Score each candidate by newly covered cells, review cost, naturalness, cluster independence, and duplication.
  7. Select iteratively until no eligible high-priority cell remains uncovered or the governing reviewers defer it.
  8. Freeze family/speaker/text-disjoint development and final sets before training.

This adopts the researcher-in-the-loop set-cover principle demonstrated by SMOLSENT without copying its English inventory or its 863-sentence count. Kuku Yalanji's final number must emerge from its own governed inventory.

Required row schema

Every retained bilingual sentence should carry at least:

{
  "id": "immutable-row-id",
  "source_text": "English source",
  "target_text": "Kuku Yalanji target",
  "origin": "speaker_authored_or_grammar_attested_or_reviewed_elicitation",
  "speaker_cluster_id": "stable-or-protected-id",
  "text_cluster_id": "stable-id",
  "source_citation": "work-page-line-or-recording-span",
  "lexeme_sense_ids": ["stable-sense-id"],
  "morphosyntactic_cells": ["documented-cell-id"],
  "variety": "documented-variety-or-unknown",
  "register": "documented-register-or-unknown",
  "rights": {
    "training": false,
    "redistribution": false,
    "hosted_provider_transfer": false,
    "model_weight_distribution": false
  },
  "review": {
    "fluent_review_status": "pending",
    "reviewer_authority": "recorded-separately",
    "review_date": null
  },
  "split_lock": "train_or_dev_or_final_or_excluded"
}

Unknown values must remain explicit. Missing provenance must not silently become approval.

Next model experiment after the data gates

The next model study should be narrower than the completed screen.

  1. Keep deterministic dictionary lookup as the production route for governed lexical queries.
  2. Use separate model-visible tasks:
<lexeme> woman
<translate> The woman returned.
<translate> The woman returned. <glossary> woman = jalbu
<morphology> [documented structured condition]
  1. Use a clean no-Bible positive objective. A frozen Bible set may remain a catastrophic-forgetting diagnostic but cannot select a checkpoint or offset failure elsewhere.
  2. Audit and train the custom gvn_Latn and task-control token rows. In the current lineage, the copied gvn_Latn embedding remained bit-identical to its initialization because embeddings were frozen.
  3. Test the collision-free tokenizer extension as a separate factor after the trainable-language-token arm, not simultaneously.
  4. Use exact max_steps, paired seeds, fixed token budgets, fixed learning-rate integrals, and per-task token accounting.
  5. Keep an untouched baseline and a token-accounted continuation control.
  6. Select a recipe on aggregate development evidence, then run multiple confirmatory seeds only if the screen passes.
  7. Keep lexical reconstruction, glossary uptake, morphology, sentence adequacy, and decoder degeneration as separate endpoints.
  8. Do not let a dictionary score authorize sentence deployment.

NLLB/Formosan tutorial disposition

The project read and audited the complete FormosanBank NLLB adaptation tutorial rather than copying its commands mechanically.

Adopted ideas:

Material correction:

Literal merging of new SentencePiece pieces into NLLB can overlap the inherited language-token ID range. The project reproduced this failure: candidate lexical IDs decoded as unrelated language tags such as kon_Latn and acm_Arab. The published tutorial tokenizer shows the same structural range risk. MobTranslate therefore built a collision-free extension that preserves every inherited token ID and appends new pieces and controls after the reserved range. That prototype passes its invariants but has not been used to select or train a model.

Reusable protocol for another Aboriginal language

This workflow is intended to transfer, while the linguistic decisions and governance do not.

  1. Governance and rights: record authority, scope, restrictions, benefit, withdrawal, hosted transfer, and weight-release conditions.
  2. Sense ledger: assign stable lexical sense/POS/variety/citation-form identifiers; never collapse by English spelling alone.
  3. Evidence inventory: separate dictionary, natural text, grammar examples, usage rows, synthetic rows, and model hypotheses.
  4. Split lock: freeze source-, speaker-, text-, and lexeme-family-disjoint development/final material before fitting.
  5. Baseline routes: benchmark deterministic lookup separately from model reconstruction and sentence translation.
  6. Full-resource census: run all candidate models over complete lexical resources, but label it a census rather than a population sample.
  7. Qualitative error taxonomy: measure exact, accepted-plus-extra, orthographic, notation, wrong-known-target, source-copy, blank, and uncategorized failures.
  8. Context join: connect every lexical sense to existing natural, grammar, usage, and synthetic contexts.
  9. Salvage first: verify attested candidates before commissioning new material.
  10. Coverage selection: build the language-specific lexical, morphosyntactic, variety, and discourse universe; use researcher-in-the-loop set cover.
  11. Speaker review: obtain or approve target sentences through the relevant fluent speakers and authorities; record uncertainty and disagreement.
  12. Tokenizer/language-token audit: verify ID preservation, fertility, trainability, tied embeddings, adapter persistence, and merged parity.
  13. Controlled screen: vary one causal factor at a time under matched updates, tokens, and learning-rate exposure.
  14. Confirmatory evaluation: use independent clusters, multiple seeds, paired analysis, and a once-opened final test.
  15. Split release gates: lookup, lexical reconstruction, and sentence generation receive independent decisions.
  16. Artifact policy: retain compact adapters, manifests, predictions, code, data order, and hashes; delete redundant merged failures and optimizer checkpoints after verified preservation.

Reproduction commands

The expanded census was run with:

WORK_ROOT=/workspace/v24-lexicon-screen \
  training/translation/run_v24_expanded_lexicon_census.sh

The evidence queue was built with:

PROG=/mnt/donto-data/donto-resources/research/translation-training/kuku-yalanji-runpod-2026-06-30
V24="$PROG/runpod/v24.0-lexicon-grounded-screen-20260715T044520Z"
EXPANDED="$V24/analysis/expanded-lexicon-census-20260715"

python3 training/translation/build_sentence_evidence_review_queue.py \
  --curated-census "$PROG/prepared/kuku-yalanji-curated-dictionary-census-20260715.eng-gvn.jsonl" \
  --context-inventory "$PROG/analysis/lexeme-context-inventory-20260715/inventory.jsonl" \
  --curated-row-analysis "$EXPANDED/analysis/curated-row-analysis.jsonl" \
  --overlap-rows "$V24/analysis/lexicon-overlap-audit-20260715/overlaps.jsonl" \
  --expected-candidate B0 --expected-candidate C0 --expected-candidate L1 \
  --expected-candidate L2 --expected-candidate L4 \
  --output "$PROG/analysis/sentence-evidence-review-queue-20260715/queue.jsonl" \
  --summary-output "$PROG/analysis/sentence-evidence-review-queue-20260715/summary.json"

Integrity and retained artifacts

Artifact SHA-256
Frozen selector d4a3623c60bcbce804e132f4ff7c422b3298ea72119a4dcdbbb6d6e896693c0b
Expanded final checksum ledger 60b127c25d18d56fb0e8b415502981d17998ee7bc43ca3a912e7a8f64569a95b
Expanded raw summary 74d716346a77a902fb006ee3cffafdb4047517ecdf825ed0bdb84c190e43beb0
Expanded curated summary b7742a914398733a8d8ce6142fb7acc380839a34379e441574dbbb815260da63
Context inventory ea1ed3c7d8bcbbc484ebc257d3c50d72134d5d4b9fec1e7a7d610e8fea898272
Review queue fca2ba770ebf67358573e8854480d76356b43f06c0ccdfc85d14c68db3b39923
Queue summary b8c18ff0ac10de92a47fb4edad44aae2fc156eca927ebd519aa45f9c833fe79c
Compact v24 artifact ledger 7f615cadb78363f2c0f7c3437cce6c7e5e7d0933b0da159b0884e15332a6214d

The compact archive retains exact arm datasets, adapters, tokenizers, manifests, parity checks, predictions, analyses, logs, and 244-file checksums. Failed merged models, optimizer states, and redundant checkpoints were not retained. The A40 pod was stopped and deleted after local hash verification; runpodctl pod list returned an empty list.

Decision table

Decision Result
Train another model immediately No
Generate another generic 20,000 synthetic pairs No
Treat L1 as a successful model No
Use L1 as evidence about task/data design Yes
Expand the benchmark from 297 prompts Completed: 1,220 + 2,724 resource censuses
Qualitatively classify failures Completed and joined to context evidence
Commission 498 new sentences automatically No
Review an upper-bound pool of 498 non-name prompts / 350 forms Yes
Validate existing XIGT and usage candidates first Yes
Use deterministic dictionary lookup Yes
Mount a custom Kuku Yalanji sentence model No

References

Final conclusion

The expanded benchmark answered the immediate question. More exposure is associated with better lexical reconstruction, but the current model does not convert generic volume into reliable source-target alignment, morphology, or sentence competence. Increasing lexical replay beyond L1 is actively harmful, and L1 does not improve the zero-context cohort.

The next useful corpus is therefore not "another N synthetic sentences." It is the smallest governed set of natural, fluent-reviewed bilingual sentences that remains after sense adjudication, attested-evidence salvage, lookup routing, and coverage selection. The review ledger now identifies the maximum candidate pool and the exact decisions needed to turn it into that set. Training should resume only when those decisions produce a checksum-frozen, split-locked data manifest.