Mi'gmaq NLLB-LoRA 1.0.0-rc1 Model Card
Model ID:
mobtranslate/migmaq-nllb-lora
Version: 1.0.0-rc1
Direction: English to Mi'gmaq (eng_Latn
-> custom mic_Latn)
Status: noncommercial research candidate; not
speaker-reviewed
Released: 2026-07-12
Operator: MobTranslate
Summary
This is a first neural English-to-Mi'gmaq research candidate trained from source-attested example sentences in the Mi'gmaq/Mi'kmaq Online Talking Dictionary. It is intended for controlled experimentation, error discovery, and community review. It is not an authoritative translator, a normative grammar, a speaker, or evidence of community approval.
The final frozen-test score is chrF++ 21.43 on 742 headword-disjoint dictionary examples. Exact match is 0/742. The model can produce Mi'gmaq-shaped text without empty-output or source-copy collapse, but held-out examples show substantial lexical substitution, omission, and argument-structure risk. Public use must therefore remain visibly labelled as a research preview.
Rights and attribution
- Training data: Mi'gmaq/Mi'kmaq Online Talking Dictionary, Creative Commons Attribution-NonCommercial 4.0 International.
- Base model:
facebook/nllb-200-distilled-600M, pinned at revisionf8d333a098d19b4fd9a8b18f94170487ad3f821d, CC BY-NC 4.0. - Model artifacts: noncommercial research use only.
- Required attribution: Mi'gmaq/Mi'kmaq Online Talking Dictionary (mikmaqonline.org), CC BY-NC 4.0.
- Commercial use is not authorized by this release.
Linguistic scope
The source collection describes its written entries as contemporary Listuguj spellings. Some source records also contain Unama'ki recordings, but no audio was used to train this text model. Source-supplied speaker codes remain provenance labels only; they were not expanded into identity, dialect, or consent claims.
Mi'gmaq is not a built-in NLLB-200 language token. This run added
mic_Latn. Its embedding was numerically initialized from
tpi_Latn because that custom-token path had already been
exercised in the MobTranslate tooling. This is only an initialization
seed. It is not a genealogical, typological, dialectal, or equivalence
claim.
Dataset
| Field | Value |
|---|---|
| Release | migmaq-online-example-parallel-v1.0.0-20260712 |
| Unique sentence pairs | 7,226 |
| Train | 5,798 |
| Validation | 686 |
| Frozen test | 742 |
| Lexical auxiliary rows | 6,733, excluded from v1 sentence training |
| Structural exclusions | 1 missing-English example |
| Split leakage violations | 0 |
| Dataset archive SHA-256 | 805c225aeebf0596fa4f892d89cf930a23a395e2b58c1c059cb47a7092f86368 |
| Download | complete dataset archive |
Connected components were formed from normalized source-entry headword, Mi'gmaq sentence, English sentence, entry ID, and bilingual-pair identity. Whole components were assigned to one split. Consequently, every validation and test source headword is absent from training. This is a deliberately difficult lexical-generalization design, not a random sentence split.
Training
| Field | Value |
|---|---|
| Architecture | NLLB-200 distilled 600M plus LoRA and trainable embeddings/output head |
| GPU | NVIDIA A40, 46,068 MiB |
| Epochs | 8 |
| Optimizer steps | 1,448 |
| Effective batch size | 32 |
| Learning rate | 1e-4 |
| LoRA | rank 32, alpha 64, dropout 0.05 |
| LoRA modules | q/k/v/out projections and feed-forward
fc1/fc2 |
| Trainable saved modules | model.shared, lm_head |
| Best checkpoint | step 1440, selected by validation chrF++ |
| Trainer runtime | 1,623.3 seconds |
| Best trainer validation chrF++ | 20.9027 |
The full trainer received train and validation files only. Its
manifest contains test_file: null. Decoding was selected on
validation before the frozen test was opened.
Final decoding
- beam size: 4;
- no-repeat n-gram size: 3;
- repetition penalty: 1.10;
- length penalty: 1.0;
- maximum new tokens: 128.
This setting was selected from three validation candidates. An unconstrained candidate produced repeated trigrams in 45/686 validation outputs and failed the predeclared degeneration gate.
Evaluation results
Validation
| Metric | Result |
|---|---|
| Rows | 686 |
| chrF++ | 21.0434 |
| chrF++ 95% bootstrap interval | 20.4083 to 21.6518 |
| SacreBLEU | 1.3867 |
| SacreBLEU 95% bootstrap interval | 0.9756 to 2.1040 |
| Exact match | 0/686 |
| Paired chrF++ gain over base proxy | +9.7344; 95% interval +9.1063 to +10.3009 |
Frozen test
| Metric | Result |
|---|---|
| Rows | 742 |
| chrF++ | 21.4298 |
| chrF++ 95% bootstrap interval | 20.8128 to 22.0332 |
| SacreBLEU | 1.7795 |
| SacreBLEU 95% bootstrap interval | 0.9308 to 2.6029 |
| Exact match | 0/742 |
| Empty outputs | 0 |
| Source copies | 0 |
| Outputs equal to any training target | 0 |
| Unique outputs | 734/742 |
| Repeated-trigram outputs | 0 |
| Mean character-length ratio | 0.9063 |
| Terminal-punctuation agreement | 98.38% |
| Apostrophe-presence agreement | 89.76% |
SacreBLEU signature:
nrefs:1|case:mixed|eff:no|tok:13a|smooth:exp|version:2.6.0.
chrF++ signature:
nrefs:1|case:mixed|eff:yes|nc:6|nw:2|space:no|version:2.6.0.
Automatic string metrics compare against one dictionary reference. They cannot determine whether a different output is grammatical, dialectally appropriate, pragmatically natural, or culturally acceptable. Conversely, a plausible-looking string can receive character overlap while assigning the wrong participants or lexical predicate.
Observed error classes
The lowest-scoring held-out rows show failures that matter beyond orthographic overlap:
| English input | Model output | Source reference | Observation |
|---|---|---|---|
| My hands are tired. | Pugwelg na mawti'g. |
Npitnn gispnegl |
Core body-part and predicate material is not recovered. |
| Are your grandfather and grandmother still alive? | Gi's gi'l na nuji'j? |
Me' gmijgamij aq gugumij mimajijig? |
Kinship and coordination content is largely absent. |
| Hold on, wait for me! | Amujpa, amujpa! |
Ge's'g, pe'l esgmali! |
Imperative content is replaced by repeated generic material. |
| When he/she sent off his/her child ... | Ta'n tujiw gesatg ugjit ugjit lpa gmu'j. |
Poqtigimateg unjann ... |
Long-clause content and participant structure are severely compressed. |
| He/she teaches his/her child how to speak Mi'gmaq ... | Mi'gme'g ugjit ugjigtug. |
Gegina'muatl unjann ta'n telinnui'simg. |
Teaching, child, and manner-of-speaking content is not preserved. |
These are reference-based observations, not speaker judgements or complete morpheme-by-morpheme analyses. Qualified Mi'gmaq reviewers must assess lexical choice, person and number, animacy, transitivity, negation, word order, dialect, and pragmatics before any broader release.
Intended use
- research demonstrations with an unavoidable “not speaker-reviewed” notice;
- collection of structured error reports from qualified reviewers;
- controlled comparison with retrieval and dictionary lookup;
- reproducibility and custom-language-token pipeline research.
Prohibited or unsupported use
- authoritative, legal, medical, emergency, ceremonial, educational-assessment, or publication-ready translation;
- representing output as community-endorsed or speaker-certified;
- commercial use;
- silently replacing dictionary entries or human translation;
- reverse Mi'gmaq-to-English translation, which was not trained or evaluated;
- claims about Unama'ki written output or other dialect coverage from this run.
Artifact integrity
| Artifact | SHA-256 |
|---|---|
| Adapter weights | d85184d614f614815910442cbc5324f440c4030cdebd6b5cff2d3d7046381c67 |
| Merged weights | 8df8467e96ba480e790951f86cdc30919cd1d441faf7d1c5c5521cef1c78eb01 |
| Complete merged-model archive | 2abe9602ee2b47cee8037b6839ce8faa052f867280591d0d917b2eea4260d591 |
| Frozen-test predictions | 5d294244e21e160a8d9e8ae35eb385a93b3608cd51df534505088778767c7566 |
| Frozen-test analysis | 2e02ba0250c38cddbc54392872c7d1bd3d0ae73291eecce46b2339a0f6b50fb7 |
The training run, failed first plumbing gate, corrected merge diagnostic, validation selection, one-shot test ledger, environment lock, resource trace, and post-close checksum manifests are retained under the program research root.
Downloads and inspection:
- complete merged model, tokenizer, run contract, predictions, and reports (TAR.GZ, 2,247,094,395 bytes);
- versioned artifact directory;
- frozen-test predictions;
- guarded live research preview.
Run and serving record
The A40 pod was active for 8,186 elapsed seconds (2.2739 hours) and was deleted after local checksum and inference verification. RunPod's final ledger posted $1.020713 for 8,191,142 billed milliseconds. The live CPU float32 service loaded the verified merged weight hash, returned healthy after warmup, completed two held-out smoke requests in 5.042 and 4.011 seconds, used no cgroup swap, and had zero restarts. Peak cgroup memory was 5,371,129,856 bytes.
Frozen evaluation used CUDA bfloat16. CPU float32 beam search selected a one-character variant for one smoke row and a different later clause for another. The archived scores therefore describe the frozen CUDA evaluation backend; byte-identical strings are not promised across numerical backends. This does not make either output speaker-validated.
Promotion ruling
Machine-integrity gate: pass. The dataset, split
controls, training, save/merge/reload path, prediction counts, and
checksums are reproducible.
Linguistic-quality gate: not passed. Zero exact matches
and the documented lexical/content failures rule out authoritative
routing.
Permitted deployment: a guarded, noncommercial research
preview with source attribution and a prominent warning.
Unrestricted deployment: blocked pending qualified
Mi'gmaq review and a materially stronger evaluation result.