Translate v2
internal proofModel evaluation bench for the Kuku Yalanji training run. This page tracks the full v8-v20 diagnostic ladder, saved outputs, downloads, and resource profiles from each RunPod run.
Where the model is after 48 hours
The training pipeline works. v20 trained the full candidate corpus and fixed the length-compression failure, but it also proved that scale alone is not enough. The hard problem is now product routing, mixture control, and faithful generalization, not GPU plumbing.
Exact known text
lookup-first
Bible references, DB usage examples, and elder-shared sentence pairs should come from the approved source table, not generation.
Bible draft
v12.0
Best current Bible-draft fallback by heldout Bible chrF; still not faithful canonical reproduction.
Usage draft
v10.0
Best current heldout DB/general usage signal. Newer balanced runs have not beaten it.
Latest research
v20.0
Full-candidate-corpus diagnostic; length compression is fixed, but faithful generalization is still unresolved.
Latest run
- Run
- v20 full candidate corpus
- Rows
- 35,394 train / 4,433 validation / 4,397 test
- Corpus shape
- 20,911 source pairs expanded into tagged tasks
- Training
- 1h52m48s trainer time, 0.98 steps/s
- RunPod cost
- about $3.51
Evaluation release
Live inference is intentionally off. Each release is reviewed through saved evaluation artifacts, exact-match counts, chrF/BLEU scores, resource usage, and downloadable model files.
- Selected release
- 0.1.0-mini-pilot
- Direction
- eng-gvn
- Saved eval rows
- 128
- Release role
- kuku_yalanji_ebible_parallel_v0.1.0
Current eval summary
Kuku Yalanji · 0.1.0-mini-pilot
- BLEU
- 0.0269
- chrF
- 4.175
- Rows
- 128.0
This page shows reproducible eval artifacts only. Approved Bible text, database examples, and elder-shared sentence pairs should remain lookup-first in the product.
Dataset
kuku_yalanji_ebible_parallel_v0.1.0
Training rows
2,048
GPU
NVIDIA A40
Artifacts
model local · adapter local
Release verdict
No verdict recorded yet.
- Budget-safe proof run on the high-confidence corpus, not a production translator.
- A40 utilization was healthy: 100% max GPU, 26.5 GiB max VRAM, and 295 W max power draw.
- Use this release to test the translate/v2 model-lab path and inference server wiring.
Downloads and artifacts
Public links resolve through the model-artifact file server. Merged models are large safetensor directories.
No public downloads are attached to this release.
Saved test outputs
Held-out examples from the saved evaluation artifacts for this model release.
Sources: ebible
- BLEU
- 0.0269
- chrF
- 4.175
- Rows
- 128.0
ebible
1CO.10.15
I speak as to wise men. Judge what I say.
men men men men men men men men men men men men men men men men men men men men men men men men men men men men men men men men men men men men men men men men men men men men men men men men men men men men men men men men men men men men men men men men men men men men men men men men men men men men men men men men men men men men men men men men men men men men men men men men men men men men men men men men men men men men men men men men men men men men men men men men men men men men men men men men men men men men men men men men men men men men men men men men men men men men men men men men men men men men men men men men men men men men men men men men men men men men men men men men men men men men men men men men men men men men men men men
Yurra binal. Yurra junkaynjaku milkabu wukurril. Yurra binal ngayu junkaynjaku balkan-balkal.
ebible
1CO.10.16
The cup of blessing which we bless, isn’t it a sharing of the blood of Christ? The bread which we break, isn’t it a sharing of the body of Christ?
baraka, yg kita syukurkan, bukankah kita bersertai dgn darah Kristus? Dan roti yg kita pecahkan, bukankah kita bersertai dgn badan Kristus?
Ngana communion nukal, ngana God thankim-bungal cupmunku. Ngana cupmun nukal, Christangka nyuluku mula dajil ngananga. Ngana mayi bread dumbarril, nukal, Christangka nyuluku bangkarr dajil ngananga. Yinyamundu ngana nyubunmal.
ebible
1CO.10.23
“All things are lawful for me,” but not all things are profitable. “All things are lawful for me,” but not all things build up.
ununununununununununununununununununununununununununununununununununununununununununununununununununununununununununununununununununununununununununununununununununununununununununununununununununununununununununununununununun nunununununununununununununununununununununununununununununununununununununununununununununununununununununununununununununununununununununununununununununun
Kanbal wadu-wadu janaku balkaway, “Ngayu wawubu, ngayu balkalkuda.” Yinya nguba yurranda milkanga kadan. Kaki yurra yala balkaway, yinya ngulkurr kari yurranka. Kaki yurra yurrangaku way dungay, God bawanji, yindu-yindu kari helpim-bunganji baja.
ebible
1CO.10.4
and all drank the same spiritual drink. For they drank of a spiritual rock that followed them, and the rock was Christ.
, 8 - 8 - 8 - 8 - 8 - 8 - 8 - 8 - 8 - 8 - 8 - 8 - 8 - 8 - 8 - 8 - 8 - 8 - 8 - 8 - 8 - 8 - 8 - 8 - 8 - 8 - 8 - 8 - 8 - 8 - 8 - 8 - 8 - 8 - 8 - 8 - 8 - 8 - 8 - 8 - 8 - 8 - 8 - 8 - 8 - 8 - 8 - 8 - 8 - 8 - 8 - 8 - 8 - 8 - 8 - 8 - 8 - 8 - 8 - 8 - 8 - 8 - 8 - 8 - 8 - 8 - 8 - 8 - 8 - 8 - 8 - 8 - 8 - 8 - 8 - 8 - 8 - 8 - 8 - 8 - 8 - 8 - 8 - 8 - 8 - 8 - 8 - 8 - 8 - 8 - 8 - 8 - 8 - 8 - 8 -
Bama wubulku wawu-wularin, bana kari, Jesus Christangka jananda bana kuljimun yungan. Bamangka wubulduku yinya bana nukanda, dandiman baja. Nyuluku yala bana kuljimun, jana wubulduku yala nukanka. Jana yala nyungundumun bana nukan, jananga wawu dandimankuda. Nyulu jananji dungan-dungan, jananin ngulkurrduku kujin-kujin.
ebible
1CO.10.7
Don’t be idolaters, as some of them were. As it is written, “The people sat down to eat and drink, and rose up to play.”
: "Ayo, jangankan kita menyembah berhala, seperti yang dilakukan oleh sebagian dari mereka". Seperti yang tertulis dalam Alkitab, "Orang-orang duduk makan dan minum, lalu bangun untuk main-main".
Kanbalda yinyarrinyangka junjuy yala ngurma idol buyay-manin, balu Godkuda. Godumu kuku yalaku: “Jana bama mayinga, kamu-kamungu nukajin, wurin-wurin. Jana kiru-karimarin, kalngar-kalngarmarin, ngurmaka junjuynku idolku.”
ebible
1CO.10.9
Let’s not test Christ, as some of them tested, and perished by the serpents.
Jesus Christ, as some of them tempted, and were destroyed by snakes.
Ngana God kari baba. Kanbalda yinyarrinyangka God baban, balu nyulu jananin bayjanka. Yamba nyulu jananin nyajin, jarba kuliji yungan. Jarbangka jananin baykan, jana wularinda.
Metrics
- train loss
- 6.295
- validation loss
- 5.773
- validation bleu
- 0.0313
- validation chrf
- 4.744
- test loss
- 5.817
- test bleu
- 0.0275
- test chrf
- 4.272
- standalone bleu
- 0.0269
- standalone chrf
- 4.175
Resource profile
- Average GPU
- 40.3%
- Max GPU
- 100%
- Mean VRAM
- n/a
- Max VRAM
- 26,525 MiB
- Max power
- 295.2 W
- Cost class
- $0.44/hr
- Estimated run cost
- n/a
Version ladder
Every material training branch from the last two days is listed here so the model line is auditable. chrF is useful for comparison, but exact approved known resources stay lookup-first.
| Version | Role | Bible ref | Usage | Elder rows | Decision | Links |
|---|---|---|---|---|---|---|
| v24.3-joint-lexeme-dose29-s3598-20260715 v24.3-joint-lexeme-dose29-s3598-20260715 | Current runtime-verified Kuku Yalanji research candidate for closed-set lexical reconstruction and separately labelled sentence drafts runtime verified research candidate | n/a | 53.17 | 0/43 | The governed training-overlapping lexical gate and every frozen development retention check passed. ADVANCE authorizes further research, not unrestricted deployment: the model remains 0/43 exact on the elder diagnostic and lacks speaker-diverse human evaluation. | |
| v21.2-claude-balanced-replay-guarded-20260714 v21.2-claude-balanced-replay-guarded-20260714 | Current Kuku Yalanji research model of record: frozen v21.2 weights with validated guarded decoding; CPU inference is intentionally unloaded research only | 43.84 | 47.96 | 0/43 | Retain the exact published v21.2 weights and guarded decoder as reproducible research artifacts, but do not load them for public CPU inference while natural elder-register transfer remains unsolved. The homepage uses separately labelled dictionary-context prompting until a stronger RunPod candidate clears the frozen linguistic and safety gates and receives appropriate speaker review. | |
| v23.0-attested-narrative-adaptation-failed v23.0-attested-narrative-adaptation-failed | Rejected three-seed attested-narrative adaptation retained as compact negative evidence negative result | n/a | n/a | n/a | Do not promote. Seed 73 improved the held-out Text 3 corpus chrF++ point estimate by 0.6414, below the preregistered +1.0 floor; its paired interval crossed zero, untagged synthetic repetition worsened, and isolated lexical reconstruction fell to 46/297. | |
| v22.0-step-matched-replay-3120-failed v22.0-step-matched-replay-3120-failed | Rejected one-variable replay-exposure experiment retained as negative evidence negative result | n/a | n/a | n/a | Do not promote. Step 3,120 failed the frozen greedy checkpoint gate because tagged synthetic repeated-segment rows were 37 against a maximum of 25, and paired comparison showed lower tagged synthetic, untagged synthetic, and dictionary-usage agreement than the final v21.2 step 4,155 checkpoint. | |
| v21.2-claude-balanced-replay-v2-candidate v21.2-claude-balanced-replay-v2-candidate | Historical v21.2 weights release; superseded for serving by the guarded-policy release without changing weights research only | 43.91 | 47.14 | 0/43 | Balanced replay preserves synthetic performance while materially recovering dictionary-usage and Bible retention relative to v21.1. Natural elder-register transfer remains unsolved, and synthetic morpheme loops remain a hard safety failure; serve only as a clearly labelled research draft pending speaker review. | |
| v21.1-codex-synthetic-direct-v2-candidate v21.1-codex-synthetic-direct-v2-candidate | Kuku Yalanji translation model v2 candidate research only | 28.76 | 34.40 | 0/43 | The model learned the synthetic treatment and remained stable without the task tag, but elder-shared, dictionary-usage, and Bible controls regressed severely. Preserve as a research artifact; keep retrieval-first routing and require balanced replay plus speaker review before promotion. | |
| v20.0-full-candidate-corpus-gvn v20.0-full-candidate-corpus-gvn | Full-candidate-corpus diagnostic artifact internal proof | 39.39 | 43.26 | 0/43 | Full-corpus v20 fixed the length-compression failure and completed the 35k-row diagnostic, but it is not a faithful general model: Bible heldout stays near 39 chrF, heldout usage is 43.26 chrF, and elder-shared sentence pairs regressed to 0/43 exact. | |
| v19.0-balanced-replay-from-v12-gvn v19.0-balanced-replay-from-v12-gvn | Balanced replay artifact internal proof | 44.29 | 53.03 | 29/43 | Restores near-v12 Bible heldout while adding elder sentence-pair/usage signal, but does not replace v10 usage or v12 Bible routing. | |
| v18.0-usage-elder-sentence-continuation-from-v10 v18.0-usage-elder-sentence-continuation-from-v10 | Elder sentence-pair memorization proof artifact internal proof | 41.25 | 54.52 | 43/43 | Exactly memorized all elder-shared sentence pairs, but regressed Bible and did not beat v10 on heldout usage. | |
| v15.0-soft-lexical-hint-bible-gvn-token v15.0-soft-lexical-hint-bible-gvn-token | Lexical-hint diagnostic negative result | 43.72 | n/a | n/a | Reproduced train rows strongly, but did not beat v12 on heldout Bible. | |
| v13.0-retrieval-context-bible-gvn-token v13.0-retrieval-context-bible-gvn-token | Retrieval-prefix diagnostic negative result | 29.71 | n/a | n/a | Retrieval context as an NLLB source prefix was harmful; heldout Bible dropped sharply. | |
| v12.0-tagged-direct-plus-reference-bible-gvn-token v12.0-tagged-direct-plus-reference-bible-gvn-token | Current Bible draft fallback route candidate | 44.36 | n/a | n/a | Best Bible draft heldout score before v19, but still zero exact heldout reproduction. | |
| v11.0-byt5-bible-control-32row v11.0-byt5-bible-control-32row | Byte-model control negative result | 5.71 | n/a | n/a | Failed the tiny overfit/control path; NLLB/LoRA remains the proven memorizing path for now. | |
| v10.0-tagged-bible-plus-glossary-usage-tpi v10.0-tagged-bible-plus-glossary-usage-tpi | Current usage/general draft fallback route candidate | 43.29 | 56.18 | n/a | Best heldout usage signal so far; weaker Bible than v12/v19. | |
| v9.8-tagged-bible-plus-db-usage-tpi v9.8-tagged-bible-plus-db-usage-tpi | First DB-usage multitask diagnostic internal proof | 44.04 | 39.47 | n/a | Memorized DB train examples but did not generalize well to word-id heldout usage. | |
| v9.7-tagged-direct-plus-reference-bible-tpi v9.7-tagged-direct-plus-reference-bible-tpi | Reference-conditioning baseline internal proof | 44.34 | n/a | n/a | Established tagged Bible direct/reference training as viable, but still zero exact heldout reproduction. | |
| v8.0-diagnostic-gates-summary v8.0-diagnostic-gates-summary | Pipeline proof gate internal proof | n/a | n/a | n/a | Proved tokenizer/LoRA/training/merge path could memorize; shifted problem from plumbing to data/task/generalization. | |
| 0.7.0-full-tpi-proxy-1.3b 0.7.0-full-tpi-proxy-1.3b | kuku_yalanji_ebible_parallel_v0.1.0 internal proof | n/a | n/a | n/a | No verdict recorded. | |
| 0.4.0-full-gvn-token 0.4.0-full-gvn-token | kuku_yalanji_ebible_parallel_v0.1.0 internal proof | n/a | n/a | n/a | No verdict recorded. | |
| 0.1.0-mini-pilot 0.1.0-mini-pilot | kuku_yalanji_ebible_parallel_v0.1.0 internal proof | n/a | n/a | n/a | No verdict recorded. | |
| 0.1.0-smoke 0.1.0-smoke | kuku_yalanji_ebible_parallel_smoke_v0.1.0 internal proof | n/a | n/a | n/a | No verdict recorded. | |
| 0.1.0-baseline 0.1.0-baseline | kuku_yalanji_ebible_parallel_v0.1.0 training ready | n/a | n/a | n/a | No verdict recorded. | |
| 1.0.0-rc1 1.0.0-rc1 | First leakage-controlled Mi'gmaq dictionary-example research candidate research only | n/a | n/a | n/a | Machine-integrity gates pass, but linguistic-quality promotion does not: frozen-test chrF++ is 21.43 with 0/742 exact matches and material lexical/content failures. Permit only a visibly warned research preview pending qualified Mi'gmaq review. |