# Kuku Yalanji RunPod Model Evaluation and Fast Feedback Loop

**Status:** completed and checksum-verified on 2026-07-14  
**Scope:** English to Kuku Yalanji (`eng_Latn` -> `gvn_Latn`) research translation  
**Benchmark:** `kuku-yalanji-lexicon-elder-version-probe-v1`  
**Final run:** `kuku-version-benchmark-runpod-a40-r2-20260714T104027Z`  
**Decision:** retain v21.2 step 4,155 with the guarded decoder; do not promote v22 step 3,120  
**Certification:** none. This is machine regression evidence, not speaker or community approval.

## 1. Purpose

This work turns model-version evaluation into a sealed, portable RunPod job. One command now:

1. verifies the kit, data, model weights, tokenizers, and configurations by SHA-256;
2. refuses the wrong Python, Torch, CUDA, GPU, or scoring-library versions;
3. enforces deterministic float32 inference with sampling and TF32 disabled;
4. evaluates each checkpoint on one frozen 340-row benchmark;
5. scores dictionary and elder rows separately;
6. computes 10,000-sample paired row-level intervals;
7. compares every GPU prediction with frozen local baselines;
8. records process and GPU resource evidence; and
9. writes `RUN_COMPLETE` only after output checksums verify.

The downloadable kit is about 100 KB compressed. It contains code, protocols, frozen data, and row-level baselines,
not multi-gigabyte weights. Checkpoints are separately versioned payloads accepted only when all registered hashes
match.

## 2. Linguistic result

The isolated-word metric accepts every canonical dictionary headword attached to the same English gloss. Elder
sentences are never pooled with isolated words.

| Checkpoint | Dictionary exact accepted | Seen target forms | Unseen target forms | Elder exact | Elder corpus chrF++ |
|---|---:|---:|---:|---:|---:|
| step 2,770 | 42/297 (14.14%) | 40/192 (20.83%) | 2/105 (1.90%) | 0/43 | 29.18 |
| step 3,120 | 43/297 (14.48%) | 41/192 (21.35%) | 2/105 (1.90%) | 0/43 | 29.26 |
| **step 4,155 guarded** | **48/297 (16.16%)** | **46/192 (23.96%)** | **2/105 (1.90%)** | **0/43** | **29.43** |

| Comparison | Dictionary exact difference | Paired 95% interval | Elder sentence chrF++ difference | Paired 95% interval |
|---|---:|---:|---:|---:|
| step 4,155 minus step 2,770 | +2.02 points | +0.67 to +3.70 | +0.27 | -0.51 to +1.22 |
| step 4,155 minus step 3,120 | +1.68 points | +0.34 to +3.37 | +0.23 | -0.42 to +0.99 |
| step 3,120 minus step 2,770 | +0.34 points | -1.01 to +1.68 | +0.04 | -0.86 to +0.84 |

The dictionary signal favors step 4,155. Every elder interval includes zero, every checkpoint is 0/43 elder exact,
and unseen-form retrieval is 2/105 for all three. Continued exposure improved retrieval of forms already represented
in training; it did not solve productive lexical generalization or natural elder-register translation.

## 3. Reproducibility

The final A40 run matched the local CPU baselines on all **1,020/1,020 predictions**:

| Checkpoint | GPU versus local | Parity |
|---|---:|---:|
| step 2,770 | 340/340 | 100% |
| step 3,120 | 340/340 | 100% |
| step 4,155 guarded | 340/340 | 100% |

An independent earlier A40 execution also matched the final A40 execution on all **1,020/1,020 predictions**. Its
result seal was rejected for a metadata-ordering bug, but its already-written predictions were preserved. This gives
both cross-device parity and same-device repeatability.

The only nonzero GPU-minus-local metric delta was about `3.55e-15` BLEU for step 2,770, a corpus-aggregation rounding
artifact. Row text, chrF++, exact match, empty counts, and copy counts were identical.

## 4. Frozen contract

### 4.1 Data

- 297 unique isolated English dictionary glosses;
- all canonical headwords for a gloss retained as accepted references;
- 43 unchanged, rights-cleared elder sentence pairs;
- 340 `eng-gvn` rows with unique IDs;
- dictionary and elder rows separated by `pair_kind`.

Input SHA-256:

```text
6e525f16eb7475c6bc5450b978d70db0fbf1a0d71f2af9f143fb80c4bf60c01f
```

The dictionary informed earlier corpus work and is not an untouched test sample. Elder rows were excluded from
training but observed in earlier evaluations. This is a version and coverage probe, not a population estimate.

### 4.2 Inference

```text
batch_size=8
max_source_length=192
max_new_tokens=208
num_beams=1
no_repeat_ngram_size=4
repetition_penalty=1.10
length_penalty=1.0
dtype=float32
do_sample=false
seed=0
deterministic_algorithms=true
allow_tf32=false
CUBLAS_WORKSPACE_CONFIG=:4096:8
NVIDIA_TF32_OVERRIDE=0
```

No test-set decoder search is allowed. Any change requires a new kit version.

### 4.3 Statistics

- 10,000 paired row-level bootstrap resamples;
- frozen seed `kuku-version-probe-2026-07-14`;
- exact retrieval on dictionary rows;
- sentence chrF++ separately for dictionary and elder rows;
- intervals describe these rows, not speakers, dialects, texts, training seeds, or future use.

## 5. Model identities

| Label | Step | Weight SHA-256 | Role |
|---|---:|---|---|
| `step2770` | 2,770 | `1f2a31593a350f373f2c4345f1bd4f16a29ece18b53fa552ba1828acff7c84f7` | earlier trajectory point |
| `step3120` | 3,120 | `8ce31236b051195fbbd01df8d7149ce729249a30959592e9762b2dabf404c2aa` | failed v22 checkpoint |
| `step4155_guarded` | 4,155 | `7f9d0fe325e9e4568e45f13179adb336b93bbd53e83ddab2826e999eba3c76f7` | current v21.2 weights |

The manifest also pins added tokens, model and generation configs, SentencePiece data, special-token maps, full
tokenizer JSON, and tokenizer config for every checkpoint. Correct weights with a changed tokenizer are rejected.

## 6. Exact RunPod runtime

```text
image=runpod/pytorch:2.4.0-py3.11-cuda12.4.1-devel-ubuntu22.04
python=3.11.10
torch=2.4.1+cu124
cuda=12.4
gpu=NVIDIA A40
transformers=4.48.3
tokenizers=0.21.4
sacrebleu=2.6.0
sentencepiece=0.2.0
safetensors=0.8.0
```

The secure-cloud pod was in `EU-SE-1`. Runtime identity is fail-closed.

## 7. One-command use

```text
kuku-eval-kit/
  benchmark_manifest.json
  requirements.lock
  SHA256SUMS.kit
  bootstrap_and_run.sh
  run_benchmark.sh
  code/
  data/
  baseline/
  models/{step2770,step3120,step4155_guarded}/
```

```bash
cd /workspace/kuku-eval-kit
./bootstrap_and_run.sh
```

The runner has no undeclared `jq`, `/usr/bin/time`, or model-service dependency. A valid result includes the frozen
manifest and lock, preflight, environment, predictions, paired analysis, parity report, resource logs,
`OUTPUT_SHA256SUMS`, and `RUN_COMPLETE`.

```bash
test -f RUN_COMPLETE
sha256sum -c OUTPUT_SHA256SUMS
```

A PID or partial output is not completion.

## 8. Faster future loop

Keep the kit and stable baselines on the training pod:

1. Train one preregistered candidate.
2. Merge it into a Transformers directory.
3. Hash every behavior-affecting file and add one unique manifest entry.
4. Run `./bootstrap_and_run.sh` on the pod.
5. Monitor prediction count, process state, and actual GPU use.
6. Download the roughly 1 MB result first and verify it locally.
7. Apply frozen linguistic, degeneration, safety, and rights gates.
8. Download or publish multi-gigabyte weights only if still promotable.
9. Preserve negative results and terminate the pod.

All three 340-row evaluations took about 3.5 minutes combined. One new checkpoint should take roughly one to two
minutes once its weights are already on the training pod.

## 9. Resources, transfer, and cost

| Checkpoint | Wall time | Peak process RSS |
|---|---:|---:|
| step 2,770 | 52.59 s | 9.71 GiB |
| step 3,120 | 76.48 s | 9.69 GiB |
| step 4,155 guarded | 81.72 s | 8.22 GiB |

Across 43 five-second GPU samples, maximum utilization was 49%, memory 6,079 MiB, power 167.2 W, and temperature
50 C. Zero outputs were empty.

Step 3,120 reused step 2,770 only as an rsync delta basis: 4,352,652,098 logical bytes, 2,216,707,248 literal bytes,
2,135,944,850 matched bytes, 1.96x speedup, and 4m29s. Final SHA-256, not similarity, established integrity.

The pod existed from 10:06:45 to 10:48:07 UTC, 0.6894 hours. At `$0.44/hour`, estimated pod cost is **$0.3033**.
The account changed by $0.3299 while other resources consumed $0.033/hour, so gross balance delta is not an isolated
bill. Pod-specific billing had not populated at deletion. The pod was deleted and the final pod list was `[]`.

## 10. Failure ledger

Each failure occurred before valid completion and improved the harness:

1. The image lacked `rsync`; it was installed only for the one-time transfer.
2. `rsync -a` returned code 23 because the network volume rejected `chown`; `-rt` omitted unsupported ownership, and hashes passed.
3. Launch 1 stopped before creating results because `jq` was absent; manifest parsing now uses pinned Python.
4. Launch 2 stopped before creating results because GNU `time` was absent; `measure_command.py` now records resources with Python's standard library.
5. The first full run appended to `run.log` after creating its checksum inventory. Independent verification rejected the seal.

For item 5, the complete failed-seal directory was preserved. The runner now closes logs, writes and verifies the
checksum inventory, and only then creates `RUN_COMPLETE`. The benchmark was rerun from scratch; remote and downloaded
checksums passed, and both A40 runs had identical predictions.

## 11. Promotion rule

This diagnostic is necessary but not sufficient. A future model must use a preregistered treatment, avoid test-set
tuning, pass empty/repetition/degeneration gates, preserve broader synthetic/usage/Bible/elder controls, report
paired differences, improve beyond seen forms, preserve lookup-first governed resources, respect rights, and receive
qualified speaker/community review before authoritative claims.

This run promotes nothing. It confirms retaining exact v21.2 step-4,155 weights with beams 1, no-repeat 4,
repetition penalty 1.10, and length penalty 1.0.

### 11.1 Current serving posture

On 2026-07-14 the operator paused all custom-model CPU inference. The Kuku Yalanji and Mi'gmaq systemd units are
stopped and disabled, the web service no longer depends on them, ports 7955 and 8766 are closed, and custom-model
endpoint variables are blank. v21.2 remains the checksummed model of record and is still downloadable, but it is not
loaded for live requests. The homepage uses the separate full-dictionary context prompt with `gpt-5.4-mini`, clearly
labelled as dictionary-guided AI requiring community verification. A custom model returns only after a stronger
RunPod candidate passes the frozen linguistic, degeneration, rights, and safety gates.

## 12. Public artifacts

- [Sealed evaluation kit](/model-artifacts/kuku-yalanji-nllb-lora/v21.2-claude-balanced-replay-guarded-20260714/lexicon-elder-benchmark-20260714/runpod-evaluation-kit-v1.tar.gz)
- [Unpacked kit](/model-artifacts/kuku-yalanji-nllb-lora/v21.2-claude-balanced-replay-guarded-20260714/lexicon-elder-benchmark-20260714/runpod-evaluation-kit-v1/)
- [A40 result bundle](/model-artifacts/kuku-yalanji-nllb-lora/v21.2-claude-balanced-replay-guarded-20260714/lexicon-elder-benchmark-20260714/kuku-yalanji-runpod-a40-benchmark-results-20260714.tar.gz)
- [A40 outputs](/model-artifacts/kuku-yalanji-nllb-lora/v21.2-claude-balanced-replay-guarded-20260714/lexicon-elder-benchmark-20260714/runpod-a40-20260714/)
- [GPU/local parity](/model-artifacts/kuku-yalanji-nllb-lora/v21.2-claude-balanced-replay-guarded-20260714/lexicon-elder-benchmark-20260714/runpod-a40-20260714/gpu-local-parity.json)
- [A40 repeatability](/model-artifacts/kuku-yalanji-nllb-lora/v21.2-claude-balanced-replay-guarded-20260714/lexicon-elder-benchmark-20260714/gpu-repeatability.json)
- [Current model API](/api/v1/models/kuku-yalanji-nllb-lora)
- [Current merged model](/model-artifacts/kuku-yalanji-nllb-lora/v21.2-claude-balanced-replay-guarded-20260714/kuku-yalanji-v21.2-guarded-merged.tar.gz)

```text
evaluation kit: 96b2a2dc0817ec20245ce0401c34c011fc0d1528a1a5cac8ec2a99ea77b8e526
OUTPUT_SHA256SUMS: e2d0fbf46c4f185a7d836700c577c380b1bfa484703793f62c6b495ac81efc28
analysis: fe4c187795d7cdd89977d8ecb6faec0cf484596f5c84c90b2582f4cbc147e950
GPU/local parity: 1434135a386825dd523e27a36ed426a612d539be16880dadb0a5199056c955aa
A40 repeatability: 363812ecbddee4b2433dab2f86bde7d66187946abdf8677757c5d9b5ff7f35c4
```

## 13. Archive, code, and limits

Canonical internal archive:

```text
/mnt/donto-data/donto-resources/research/translation-training/kuku-yalanji-runpod-2026-06-30/
  runpod/kuku-version-benchmark-runpod-20260714T100622Z/
```

Reusable code lives in `training/translation/`: `evaluate_nllb_lora.py`, `analyze_kuku_version_probe.py`,
`verify_kuku_benchmark_inputs.py`, `compare_kuku_benchmark_runs.py`, `measure_command.py`,
`run_kuku_version_benchmark.sh`, and `bootstrap_kuku_version_benchmark.sh`.

Automatic overlap and headword retrieval do not establish grammaticality, adequacy, discourse naturalness, dialect
fit, cultural acceptability, pronunciation, or speaker endorsement. The next stronger evaluation needs untouched
texts, blinded speaker ratings, explicit morphology/agreement contrasts, unseen-lemma composition, linguistic error
annotation, multiple training seeds, and separate forward/reverse directions. Until then, this is a guarded,
lookup-first, noncommercial research preview.
