Skip to main content

Translate v2

research only

Model evaluation bench for the Kuku Yalanji training run. This page tracks the full v8-v20 diagnostic ladder, saved outputs, downloads, and resource profiles from each RunPod run.

Where the model is after 48 hours

The training pipeline works. v20 trained the full candidate corpus and fixed the length-compression failure, but it also proved that scale alone is not enough. The hard problem is now product routing, mixture control, and faithful generalization, not GPU plumbing.

Saved evals only

Exact known text

lookup-first

Bible references, DB usage examples, and elder-shared sentence pairs should come from the approved source table, not generation.

Bible draft

v12.0

Best current Bible-draft fallback by heldout Bible chrF; still not faithful canonical reproduction.

Usage draft

v10.0

Best current heldout DB/general usage signal. Newer balanced runs have not beaten it.

Latest research

v20.0

Full-candidate-corpus diagnostic; length compression is fixed, but faithful generalization is still unresolved.

Latest run

Run
v20 full candidate corpus
Rows
35,394 train / 4,433 validation / 4,397 test
Corpus shape
20,911 source pairs expanded into tagged tasks
Training
1h52m48s trainer time, 0.98 steps/s
RunPod cost
about $3.51

Evaluation release

Live inference is intentionally off. Each release is reviewed through saved evaluation artifacts, exact-match counts, chrF/BLEU scores, resource usage, and downloadable model files.

Selected release
v21.2-claude-balanced-replay-guarded-20260714
Direction
eng-gvn
Saved eval rows
see artifacts
Release role
Current Kuku Yalanji research model of record: frozen v21.2 weights with validated guarded decoding; CPU inference is intentionally unloaded

Current eval summary

Kuku Yalanji · v21.2-claude-balanced-replay-guarded-20260714

No compact saved-output preview is attached to this release yet. Use the artifacts and metrics sections below.

This page shows reproducible eval artifacts only. Approved Bible text, database examples, and elder-shared sentence pairs should remain lookup-first in the product.

Dataset

Byte-identical v21.2 balanced-replay weights trained on 22,164 rows per epoch; this release changes decoding only after a separately frozen transfer protocol

Training rows

planned

GPU

NVIDIA A40 (v22 experiment and decoder-transfer validation)

Artifacts

model public · adapter n/a

Release verdict

Retain the exact published v21.2 weights and guarded decoder as reproducible research artifacts, but do not load them for public CPU inference while natural elder-register transfer remains unsolved. The homepage uses separately labelled dictionary-context prompting until a stronger RunPod candidate clears the frozen linguistic and safety gates and receives appropriate speaker review.

  • The merged model SHA-256 is exactly the published v21.2 identity; this is a decoding-policy release, not a renamed checkpoint.
  • The decoder was selected on the v22 development set and locked before v22 test/control inference and before any v21.2 transfer inference.
  • A first transfer attempt reconstructed non-identical float32 bytes and was rejected before evaluation; only the exact published v21.2 merged model was scored.
  • Guarded versus greedy tagged synthetic corpus chrF changed by +1.8723 while repeated-segment rows fell from 37 to 2; all seven frozen sets had zero empty outputs.
  • Exact match fell on some sets and Bible chrF changed slightly downward; these outcomes are reported in the experiment book rather than hidden.
  • A post-training 297-prompt canonical-headword probe scored this release at 48/297 normalized accepted-reference exact matches (16.16%); the result was 46/192 for target forms seen as training tokens and 2/105 for unseen targets.
  • Under the same fixed decoder, step 4,155 exceeded step 3,120 by 1.68 exact-match percentage points on the lexicon probe (paired 95% interval +0.34 to +3.37), while the 43-row elder difference remained inconclusive and every checkpoint was 0/43 exact.
  • An exact NVIDIA A40 reproduction evaluated all 340 frozen rows for steps 2,770, 3,120, and 4,155. All 1,020 GPU predictions matched the frozen local baselines, and a second A40 run reproduced all 1,020 predictions exactly.
  • The portable evaluator fail-closes on input, model-file, runtime, decoder, determinism, GPU-use, empty-output, baseline-parity, and output-checksum violations. Its compressed kit is 100,526 bytes and excludes model weights.
  • A first complete A40 attempt was rejected when a post-seal log append invalidated its checksum inventory. The failed attempt and repair evidence remain in the private research archive; only the clean from-scratch rerun is published.
  • On 2026-07-14 the Kuku Yalanji and Mi'gmaq CPU inference workers were stopped and disabled, their web-service dependencies were removed, and custom-model endpoints were cleared. This release remains downloadable and reproducible but is not currently loaded for live translation.
  • Automatic metrics are regression evidence, not community validation. Lookup-first routing and explicit research-draft labelling remain required.

Downloads and artifacts

Public links resolve through the model-artifact file server. Merged models are large safetensor directories.

Comprehensive model, training, and hosting guidedocumentation · htmlv22 step-matched replay and decoder-transfer experimentdocumentation · htmlRunPod evaluation and fast feedback loopdocumentation · htmlRunPod evaluation and fast-loop sourcedocumentation · markdownPortable RunPod evaluation kitevaluation · tar.gzPortable RunPod kit instructionsdocumentation · markdownSealed A40 benchmark result bundleevaluation · tar.gzSealed A40 output checksumsmetadata · sha256A40 version-probe analysisevaluation · jsonA40 versus local prediction parityevaluation · jsonA40 repeatability comparisonevaluation · jsonA40 resource summaryevaluation · jsonLexicon and elder version-probe resultsevaluation · markdownLexicon and elder version-probe analysisevaluation · jsonFrozen lexicon and elder benchmark inputdataset · jsonlLexicon-probe manifestmetadata · jsonLexicon and elder benchmark preregistrationevaluation · markdownLexicon and elder benchmark run contractevaluation · markdownLexicon and elder benchmark package checksumsmetadata · sha256Step 2,770 row-level lexicon and elder predictionsevaluation · jsonStep 3,120 row-level lexicon and elder predictionsevaluation · jsonStep 4,155 guarded row-level lexicon and elder predictionsevaluation · jsonAtlas dynamic LoRA guarded hosting bundlebundle · tar.gzStandalone guarded merged-model bundlemodel · tar.gzGuarded hosting manifestmetadata · jsonRelease metadatametadata · jsonBundle READMEdocumentation · markdownRelease checksumsmetadata · sha256Local release verificationevaluation · textFrozen decoder-transfer protocolevaluation · markdownDecoder-transfer promotion gateevaluation · jsonExact v21.2 training splitsdataset · tar.gzComplete synthetic process corpus v2dataset · zip

Metrics

weights train loss
1.199
synthetic dev chrf
55.2
synthetic dev exact
394.0
synthetic dev segment loop rows
1.000
synthetic test tagged chrf
55.7
synthetic test tagged exact
426.0
synthetic test tagged segment loop rows
2.000
synthetic test untagged chrf
55.7
synthetic test untagged exact
422.0
synthetic test untagged segment loop rows
2.000
elder shared chrf
29.4
elder shared exact
0.0000
heldout usage chrf
48.0
heldout usage exact
1.000
bible direct chrf
43.7
bible ref chrf
43.8
empty outputs all sets
0.0000
isolated lexicon probe rows
297.0
isolated lexicon probe normalized exact
48.0
isolated lexicon probe normalized exact percent
16.2
isolated lexicon probe mean sentence chrf
30.4
isolated lexicon probe seen target exact percent
24.0
isolated lexicon probe unseen target exact percent
1.905
elder version probe chrf
29.4
elder version probe exact
0.0000
runpod a40 benchmark rows
340.0
runpod a40 predictions evaluated
1020.0
runpod a40 gpu local identical predictions
1020.0
runpod a40 repeat identical predictions
1020.0
runpod a40 empty outputs
0.0000
runpod a40 total inference seconds
210.8
runpod a40 max gpu memory mib
6079.0
num beams
1.000
no repeat ngram size
4.000
repetition penalty
1.100
length penalty
1.000
merged model sha256
7f9d0fe325e9e4568e45f13179adb336b93bbd53e83ddab2826e999eba3c76f7

Resource profile

Average GPU
n/a
Max GPU
n/a
Mean VRAM
n/a
Max VRAM
n/a
Max power
n/a
Cost class
$0.44/hr
Estimated run cost
$0.8008

Version ladder

Every material training branch from the last two days is listed here so the model line is auditable. chrF is useful for comparison, but exact approved known resources stay lookup-first.

23 releases
VersionRoleBible refUsageElder rowsDecisionLinks
v24.3-joint-lexeme-dose29-s3598-20260715

v24.3-joint-lexeme-dose29-s3598-20260715

Current runtime-verified Kuku Yalanji research candidate for closed-set lexical reconstruction and separately labelled sentence drafts

runtime verified research candidate

n/a53.170/43

The governed training-overlapping lexical gate and every frozen development retention check passed. ADVANCE authorizes further research, not unrestricted deployment: the model remains 0/43 exact on the elder diagnostic and lacks speaker-diverse human evaluation.

v21.2-claude-balanced-replay-guarded-20260714

v21.2-claude-balanced-replay-guarded-20260714

Current Kuku Yalanji research model of record: frozen v21.2 weights with validated guarded decoding; CPU inference is intentionally unloaded

research only

43.8447.960/43

Retain the exact published v21.2 weights and guarded decoder as reproducible research artifacts, but do not load them for public CPU inference while natural elder-register transfer remains unsolved. The homepage uses separately labelled dictionary-context prompting until a stronger RunPod candidate clears the frozen linguistic and safety gates and receives appropriate speaker review.

v23.0-attested-narrative-adaptation-failed

v23.0-attested-narrative-adaptation-failed

Rejected three-seed attested-narrative adaptation retained as compact negative evidence

negative result

n/an/an/a

Do not promote. Seed 73 improved the held-out Text 3 corpus chrF++ point estimate by 0.6414, below the preregistered +1.0 floor; its paired interval crossed zero, untagged synthetic repetition worsened, and isolated lexical reconstruction fell to 46/297.

v22.0-step-matched-replay-3120-failed

v22.0-step-matched-replay-3120-failed

Rejected one-variable replay-exposure experiment retained as negative evidence

negative result

n/an/an/a

Do not promote. Step 3,120 failed the frozen greedy checkpoint gate because tagged synthetic repeated-segment rows were 37 against a maximum of 25, and paired comparison showed lower tagged synthetic, untagged synthetic, and dictionary-usage agreement than the final v21.2 step 4,155 checkpoint.

v21.2-claude-balanced-replay-v2-candidate

v21.2-claude-balanced-replay-v2-candidate

Historical v21.2 weights release; superseded for serving by the guarded-policy release without changing weights

research only

43.9147.140/43

Balanced replay preserves synthetic performance while materially recovering dictionary-usage and Bible retention relative to v21.1. Natural elder-register transfer remains unsolved, and synthetic morpheme loops remain a hard safety failure; serve only as a clearly labelled research draft pending speaker review.

v21.1-codex-synthetic-direct-v2-candidate

v21.1-codex-synthetic-direct-v2-candidate

Kuku Yalanji translation model v2 candidate

research only

28.7634.400/43

The model learned the synthetic treatment and remained stable without the task tag, but elder-shared, dictionary-usage, and Bible controls regressed severely. Preserve as a research artifact; keep retrieval-first routing and require balanced replay plus speaker review before promotion.

v20.0-full-candidate-corpus-gvn

v20.0-full-candidate-corpus-gvn

Full-candidate-corpus diagnostic artifact

internal proof

39.3943.260/43

Full-corpus v20 fixed the length-compression failure and completed the 35k-row diagnostic, but it is not a faithful general model: Bible heldout stays near 39 chrF, heldout usage is 43.26 chrF, and elder-shared sentence pairs regressed to 0/43 exact.

v19.0-balanced-replay-from-v12-gvn

v19.0-balanced-replay-from-v12-gvn

Balanced replay artifact

internal proof

44.2953.0329/43

Restores near-v12 Bible heldout while adding elder sentence-pair/usage signal, but does not replace v10 usage or v12 Bible routing.

v18.0-usage-elder-sentence-continuation-from-v10

v18.0-usage-elder-sentence-continuation-from-v10

Elder sentence-pair memorization proof artifact

internal proof

41.2554.5243/43

Exactly memorized all elder-shared sentence pairs, but regressed Bible and did not beat v10 on heldout usage.

v15.0-soft-lexical-hint-bible-gvn-token

v15.0-soft-lexical-hint-bible-gvn-token

Lexical-hint diagnostic

negative result

43.72n/an/a

Reproduced train rows strongly, but did not beat v12 on heldout Bible.

v13.0-retrieval-context-bible-gvn-token

v13.0-retrieval-context-bible-gvn-token

Retrieval-prefix diagnostic

negative result

29.71n/an/a

Retrieval context as an NLLB source prefix was harmful; heldout Bible dropped sharply.

v12.0-tagged-direct-plus-reference-bible-gvn-token

v12.0-tagged-direct-plus-reference-bible-gvn-token

Current Bible draft fallback

route candidate

44.36n/an/a

Best Bible draft heldout score before v19, but still zero exact heldout reproduction.

v11.0-byt5-bible-control-32row

v11.0-byt5-bible-control-32row

Byte-model control

negative result

5.71n/an/a

Failed the tiny overfit/control path; NLLB/LoRA remains the proven memorizing path for now.

v10.0-tagged-bible-plus-glossary-usage-tpi

v10.0-tagged-bible-plus-glossary-usage-tpi

Current usage/general draft fallback

route candidate

43.2956.18n/a

Best heldout usage signal so far; weaker Bible than v12/v19.

v9.8-tagged-bible-plus-db-usage-tpi

v9.8-tagged-bible-plus-db-usage-tpi

First DB-usage multitask diagnostic

internal proof

44.0439.47n/a

Memorized DB train examples but did not generalize well to word-id heldout usage.

v9.7-tagged-direct-plus-reference-bible-tpi

v9.7-tagged-direct-plus-reference-bible-tpi

Reference-conditioning baseline

internal proof

44.34n/an/a

Established tagged Bible direct/reference training as viable, but still zero exact heldout reproduction.

v8.0-diagnostic-gates-summary

v8.0-diagnostic-gates-summary

Pipeline proof gate

internal proof

n/an/an/a

Proved tokenizer/LoRA/training/merge path could memorize; shifted problem from plumbing to data/task/generalization.

0.7.0-full-tpi-proxy-1.3b

0.7.0-full-tpi-proxy-1.3b

kuku_yalanji_ebible_parallel_v0.1.0

internal proof

n/an/an/a

No verdict recorded.

0.4.0-full-gvn-token

0.4.0-full-gvn-token

kuku_yalanji_ebible_parallel_v0.1.0

internal proof

n/an/an/a

No verdict recorded.

0.1.0-mini-pilot

0.1.0-mini-pilot

kuku_yalanji_ebible_parallel_v0.1.0

internal proof

n/an/an/a

No verdict recorded.

0.1.0-smoke

0.1.0-smoke

kuku_yalanji_ebible_parallel_smoke_v0.1.0

internal proof

n/an/an/a

No verdict recorded.

0.1.0-baseline

0.1.0-baseline

kuku_yalanji_ebible_parallel_v0.1.0

training ready

n/an/an/a

No verdict recorded.

1.0.0-rc1

1.0.0-rc1

First leakage-controlled Mi'gmaq dictionary-example research candidate

research only

n/an/an/a

Machine-integrity gates pass, but linguistic-quality promotion does not: frozen-test chrF++ is 21.43 with 0/742 exact matches and material lexical/content failures. Permit only a visibly warned research preview pending qualified Mi'gmaq review.