mobtranslate.com / docs › Mi'gmaq source: migmaq-model-card.md

Mi'gmaq NLLB-LoRA 1.0.0-rc1 Model Card

Model ID: mobtranslate/migmaq-nllb-lora
Version: 1.0.0-rc1
Direction: English to Mi'gmaq (eng_Latn -> custom mic_Latn)
Status: noncommercial research candidate; not speaker-reviewed
Released: 2026-07-12
Operator: MobTranslate

Summary

This is a first neural English-to-Mi'gmaq research candidate trained from source-attested example sentences in the Mi'gmaq/Mi'kmaq Online Talking Dictionary. It is intended for controlled experimentation, error discovery, and community review. It is not an authoritative translator, a normative grammar, a speaker, or evidence of community approval.

The final frozen-test score is chrF++ 21.43 on 742 headword-disjoint dictionary examples. Exact match is 0/742. The model can produce Mi'gmaq-shaped text without empty-output or source-copy collapse, but held-out examples show substantial lexical substitution, omission, and argument-structure risk. Public use must therefore remain visibly labelled as a research preview.

Rights and attribution

Linguistic scope

The source collection describes its written entries as contemporary Listuguj spellings. Some source records also contain Unama'ki recordings, but no audio was used to train this text model. Source-supplied speaker codes remain provenance labels only; they were not expanded into identity, dialect, or consent claims.

Mi'gmaq is not a built-in NLLB-200 language token. This run added mic_Latn. Its embedding was numerically initialized from tpi_Latn because that custom-token path had already been exercised in the MobTranslate tooling. This is only an initialization seed. It is not a genealogical, typological, dialectal, or equivalence claim.

Dataset

Field Value
Release migmaq-online-example-parallel-v1.0.0-20260712
Unique sentence pairs 7,226
Train 5,798
Validation 686
Frozen test 742
Lexical auxiliary rows 6,733, excluded from v1 sentence training
Structural exclusions 1 missing-English example
Split leakage violations 0
Dataset archive SHA-256 805c225aeebf0596fa4f892d89cf930a23a395e2b58c1c059cb47a7092f86368
Download complete dataset archive

Connected components were formed from normalized source-entry headword, Mi'gmaq sentence, English sentence, entry ID, and bilingual-pair identity. Whole components were assigned to one split. Consequently, every validation and test source headword is absent from training. This is a deliberately difficult lexical-generalization design, not a random sentence split.

Training

Field Value
Architecture NLLB-200 distilled 600M plus LoRA and trainable embeddings/output head
GPU NVIDIA A40, 46,068 MiB
Epochs 8
Optimizer steps 1,448
Effective batch size 32
Learning rate 1e-4
LoRA rank 32, alpha 64, dropout 0.05
LoRA modules q/k/v/out projections and feed-forward fc1/fc2
Trainable saved modules model.shared, lm_head
Best checkpoint step 1440, selected by validation chrF++
Trainer runtime 1,623.3 seconds
Best trainer validation chrF++ 20.9027

The full trainer received train and validation files only. Its manifest contains test_file: null. Decoding was selected on validation before the frozen test was opened.

Final decoding

This setting was selected from three validation candidates. An unconstrained candidate produced repeated trigrams in 45/686 validation outputs and failed the predeclared degeneration gate.

Evaluation results

Validation

Metric Result
Rows 686
chrF++ 21.0434
chrF++ 95% bootstrap interval 20.4083 to 21.6518
SacreBLEU 1.3867
SacreBLEU 95% bootstrap interval 0.9756 to 2.1040
Exact match 0/686
Paired chrF++ gain over base proxy +9.7344; 95% interval +9.1063 to +10.3009

Frozen test

Metric Result
Rows 742
chrF++ 21.4298
chrF++ 95% bootstrap interval 20.8128 to 22.0332
SacreBLEU 1.7795
SacreBLEU 95% bootstrap interval 0.9308 to 2.6029
Exact match 0/742
Empty outputs 0
Source copies 0
Outputs equal to any training target 0
Unique outputs 734/742
Repeated-trigram outputs 0
Mean character-length ratio 0.9063
Terminal-punctuation agreement 98.38%
Apostrophe-presence agreement 89.76%

SacreBLEU signature: nrefs:1|case:mixed|eff:no|tok:13a|smooth:exp|version:2.6.0.
chrF++ signature: nrefs:1|case:mixed|eff:yes|nc:6|nw:2|space:no|version:2.6.0.

Automatic string metrics compare against one dictionary reference. They cannot determine whether a different output is grammatical, dialectally appropriate, pragmatically natural, or culturally acceptable. Conversely, a plausible-looking string can receive character overlap while assigning the wrong participants or lexical predicate.

Observed error classes

The lowest-scoring held-out rows show failures that matter beyond orthographic overlap:

English input Model output Source reference Observation
My hands are tired. Pugwelg na mawti'g. Npitnn gispnegl Core body-part and predicate material is not recovered.
Are your grandfather and grandmother still alive? Gi's gi'l na nuji'j? Me' gmijgamij aq gugumij mimajijig? Kinship and coordination content is largely absent.
Hold on, wait for me! Amujpa, amujpa! Ge's'g, pe'l esgmali! Imperative content is replaced by repeated generic material.
When he/she sent off his/her child ... Ta'n tujiw gesatg ugjit ugjit lpa gmu'j. Poqtigimateg unjann ... Long-clause content and participant structure are severely compressed.
He/she teaches his/her child how to speak Mi'gmaq ... Mi'gme'g ugjit ugjigtug. Gegina'muatl unjann ta'n telinnui'simg. Teaching, child, and manner-of-speaking content is not preserved.

These are reference-based observations, not speaker judgements or complete morpheme-by-morpheme analyses. Qualified Mi'gmaq reviewers must assess lexical choice, person and number, animacy, transitivity, negation, word order, dialect, and pragmatics before any broader release.

Intended use

Prohibited or unsupported use

Artifact integrity

Artifact SHA-256
Adapter weights d85184d614f614815910442cbc5324f440c4030cdebd6b5cff2d3d7046381c67
Merged weights 8df8467e96ba480e790951f86cdc30919cd1d441faf7d1c5c5521cef1c78eb01
Complete merged-model archive 2abe9602ee2b47cee8037b6839ce8faa052f867280591d0d917b2eea4260d591
Frozen-test predictions 5d294244e21e160a8d9e8ae35eb385a93b3608cd51df534505088778767c7566
Frozen-test analysis 2e02ba0250c38cddbc54392872c7d1bd3d0ae73291eecce46b2339a0f6b50fb7

The training run, failed first plumbing gate, corrected merge diagnostic, validation selection, one-shot test ledger, environment lock, resource trace, and post-close checksum manifests are retained under the program research root.

Downloads and inspection:

Run and serving record

The A40 pod was active for 8,186 elapsed seconds (2.2739 hours) and was deleted after local checksum and inference verification. RunPod's final ledger posted $1.020713 for 8,191,142 billed milliseconds. The live CPU float32 service loaded the verified merged weight hash, returned healthy after warmup, completed two held-out smoke requests in 5.042 and 4.011 seconds, used no cgroup swap, and had zero restarts. Peak cgroup memory was 5,371,129,856 bytes.

Frozen evaluation used CUDA bfloat16. CPU float32 beam search selected a one-character variant for one smoke row and a different later clause for another. The archived scores therefore describe the frozen CUDA evaluation backend; byte-identical strings are not promised across numerical backends. This does not make either output speaker-validated.

Promotion ruling

Machine-integrity gate: pass. The dataset, split controls, training, save/merge/reload path, prediction counts, and checksums are reproducible.
Linguistic-quality gate: not passed. Zero exact matches and the documented lexical/content failures rule out authoritative routing.
Permitted deployment: a guarded, noncommercial research preview with source attribution and a prominent warning.
Unrestricted deployment: blocked pending qualified Mi'gmaq review and a materially stronger evaluation result.