mobtranslate.com › docslanguage data and model programs

MobTranslate Documentation

Reference materials for the 20,047-pair Kuku Yalanji (gvn) English ↔ Kuku Yalanji parallel corpus — a versioned, process-audited synthetic research corpus authored against Elisabeth Patz's reference grammar and the project dictionary. Its rows are project-reviewed but remain pending Kuku Yalanji speaker review; process verification is not elder or community certification.

Model status: v24.3 is the current runtime-verified Kuku Yalanji research candidate. Under its explicit <lexeme> task it reconstructed 2,696/2,724 accepted dictionary records exactly on a governed, training-overlapping closed set, while passing every frozen development sentence-retention check. This is not unseen lexical generalisation or 98.97% sentence accuracy: the elder diagnostic remains 0/43 exact, the run used one seed and development data, and it is not speaker-certified or authoritative. The exact adapter, base, training data, complete failure ledger, CPU resource measurements, and hosting guide are downloadable; high-throughput CPU hosting is not recommended.

Evaluation loop: the frozen 340-row lexicon and elder benchmark is now a checksum-sealed, one-command RunPod kit. Two deterministic A40 executions matched each other and the local CPU baselines on all 1,020 predictions; step 4,155 remains ahead on canonical dictionary retrieval while elder differences remain inconclusive.

Model distribution: the versioned model API lists Kuku Yalanji and Mi'gmaq releases, absolute model and dataset downloads, checksums, rights, training summaries, and evaluation metadata. Its OpenAPI document is public and requires no API key.

Hybrid benchmark status: on the 43-row blind elder-shared control, Codex post-editing had the highest automatic-overlap point score, but its difference from Claude did not cleanly exclude zero. Neither route is speaker-reviewed, so LLM post-editing remains disabled while v21.2 is exposed only as a guarded research draft.

Mi'gmaq status: NLLB-LoRA 1.0.0-rc1 is live only as a guarded, noncommercial research preview. Its leakage-controlled frozen test and machine-integrity checks are complete, but exact match is 0/742 and material lexical/content failures remain. It is not speaker-reviewed or approved for authoritative translation.

The documents

DocumentWhat it is
Operator's Guide & Handoff Manual The complete working manual for how the corpus is built — the method, the rules, phantom-lexeme discipline, the batch loop, the elder-surface phonology, and every recurring pitfall. .md
Sample Production Pages What the finished products look like: one polished page each of the Teaching Grammar, the Learner's Dictionary, and the Interlinear Reader, built from verified corpus sentences. .md
Mega Grammar Cheatsheet The living grammar book: 80+ numbered rules covering the case system, verb morphology, subordination, questions, kinship, and the lexical fields, with worked examples. .md
Dictionary Errata & Phantom Watchlist The living dictionary rulings: verified lexemes, wrong-sense corrections, and the watchlist of coinages ruled out — the record that keeps the corpus honest. .md
Kuku Yalanji v21.2 Model, Training & Hosting Guide The public end-to-end release handoff: model API, immutable downloads, checksums, exact lineage, training data, RunPod procedure, Atlas dynamic LoRA and merged loading, evaluation, failure modes, rights, and update protocol. .md
Kuku Yalanji v24.3 Model, Training, Evaluation & Hosting Guide The complete working-model handoff: immutable base and adapter bundles, exact 115,136-row training schedule, two-task API, all 28 lexical failures, sentence-retention limits, measured CPU/GPU resources, RunPod recipe, multi-language adapter design, rights, and host acceptance checks. .md
v22 Step-Matched Replay & Decoder Transfer The preregistered RunPod experiment record: exact step curve, failed v22 checkpoint gate, development-only decoder selection, exact-model transfer, paired bootstrap intervals, linguistic limits, resource use, cost, hashes, and deletion proof. .md
v23 Attested-Narrative Preregistration The treatment frozen before training: research question, evidential hierarchy, speaker-disjoint splits, narrative/XIGT decisions, leakage rules, exact LoRA recipe, test lock, all-or-nothing promotion gates, statistical interpretation, and input hashes. .md
v23 Attested-Narrative Adaptation The complete three-seed negative-result record: attested split construction, leakage controls, frozen training treatment, speaker-disjoint natural test, paired intervals, dictionary CPU gate, linguistic limits, rejected-weight deletion, integrity correction, cost, and next-experiment design. .md
RunPod Evaluation & Fast Feedback Loop The portable benchmark manual and completed A40 audit: frozen linguistic protocol, exact runtime and hashes, one-command execution, CPU/GPU parity, repeatability, resources, costs, failure ledger, promotion gates, and the faster post-training loop. .md
v24 Independent-Review Disposition The decision record for the external review: what was adopted, modified, deferred, or rejected; the revised research tasks; governance and release boundary; split deployment gates; and explicit non-actions. .md
v24 Data Requirements & Sentence Commissioning The expanded 3,944-row lexical census, qualitative failure analysis, evidence-linked review queue, bounded new-context inventory, and reusable coverage-driven protocol for deciding what bilingual sentences must exist before another training run. .md
Technical Audit & v24 Training Plan The post-review technical dossier: complete data and model history, exact failures and retained results, current routing, split deployment gates, code excerpts, forensic diagnosis, and the revised preregistration-ready RunPod plan. .md
Next 16 Build Incident & Deployment Runbook The evidence-led build and deployment record: forced-Webpack diagnosis, Tailwind source discovery, output-trace repair, disk-contention measurements, isolated build procedure, atomic rollback, service gates, live API checks, and browser verification. .md
v21 Independent Model Comparison A paired, checksum-audited comparison of the Codex synthetic-only model and Claude balanced-replay model, including confidence intervals, linguistic slices, degeneration checks, limitations, and the promotion ruling. .md
Translation Systems Research & Benchmark The comprehensive custom-model, Codex, and Claude blind comparison: leakage controls, failures, provider usage, product routing, linguistic evaluation design, RunPod improvement plan, Mi'gmaq separation rules, artifact hashes, and promotion checklist. .md
Mi'gmaq Recording Integration The reproducible archive and mapping manual for the Mi'gmaq Online recordings: source rights, checksummed manifests, exact-match policy, duplicate handling, database provenance, verification gates, and recovery procedure. .md
Mi'gmaq Model Training & Handoff The living, reproducible record for the Mi'gmaq RunPod model: corpus freeze, rights and dialect scope, leakage controls, model design, mandatory evaluation gates, commands, artifacts, and promotion decisions. .md
Mi'gmaq NLLB-LoRA Model Card The user-facing card for version 1.0.0-rc1: intended use, rights, linguistic scope, training data, held-out results, observed errors, artifact hashes, and the guarded-release ruling. .md
Mi'gmaq v1 Evaluation & Promotion Report The full leakage-controlled evaluation: training curve, validation-only decoding selection, bootstrap intervals, degeneration controls, manual error analysis, resource evidence, and promotion decision. .md
Mi'gmaq Dictionary Browse Consolidation The non-destructive audit and consolidation of duplicate historical imports: pre-migration inventories, canonical identity rule, copied child content, postconditions, live count, and recovery. .md
Mi'gmaq v1 Deployment Verification The end-to-end live verification record for the translator, API guards, model and dataset downloads, responsive UI, dictionary recordings, service resources, Kuku regression check, and known external-script observation. .md

These pages are generated from the programs' own source documents and are rerendered as the corpora, audits, and model runs progress.