Skip to main content
Back to the atlas

Methods, sources & open data

How this atlas is built, and how to cite it

The Atlas of Australian Languages is assembled from open scholarly datasets and served as versioned static artifacts — never a live database query at request time. Every field on every profile carries its source, coverage is reported as honest fractions rather than rounded to imply completeness, and the whole joined dataset is downloadable below under the upstream licences. Data release 1.1.0 · build 2026-07-12.

Upstream datasets

Each source is used for a specific plane of the atlas. Versions and licences are exactly as released; where a versioned Zenodo DOI is best obtained from the dataset's own landing page we link the canonical page rather than assert an identifier we have not verified.

Upstream datasets with version, licence, canonical link or DOI, and what the atlas uses each for.
DatasetVersionLicenceLink / DOIUsed for
Glottolog5.3 CLDFCC-BY-4.0Classification chains, coordinates, ISO/glottocodes, AES endangerment.
Grambankv1.0.3 CLDFCC-BY-4.0Grammatical feature profiles (the /atlas/grammar lens + grammar-values.csv).
WALS2020.4 CLDFCC-BY-4.0Supplementary grammatical features baked into the feature tables.
Australianist typology extension (AUS extension)mobtranslate-pg typology planeCC-BY-4.0 (derived)Australia-specific features + recorded-agreement similarity neighbours.
AIATSIS AUSTLANGdata.gov.au exportCC-BY-4.0Canonical AU language codes, approximate coordinates, alt-name lists.
Bouckaert, Bowern & Atkinson 2018 (Phlorest)phlorest CLDFCC-BY-4.0The dated Pama-Nyungan phylogeographic tree (deep-time positions on /atlas/spread).
PHOIBLEreferencedCC-BY-SA-3.0Phonological inventories — referenced pointer only, not yet ingested.
E. M. Curr, The Australian Race (1886-87)archive.org OCRPublic Domain19th-century locality wordlists (historical appendix — colonial OCR, not community lexicons).
English Wiktionarykaikki.org exportCC-BY-SA-4.0Open lexical resources for some languoids.
mobtranslate-pg (live DB)read-only snapshot @ 2026-07-12T00:00:00Zmixed (see per-source)Live, community-curated dictionary word-counts (read-only build-time snapshot).

CC-BY vs CC-BY-NC.The Bouckaert et al. 2018 phylogeny is used here via the Phlorest CLDF release (CC-BY-4.0). The same tree is also distributed through D-PLACE, parts of which carry CC-BY-NC. The atlas deliberately sources the CC-BY-4.0 Phlorest release so the deep-time layer is free of a non-commercial clause. Share-alike sources — Wiktionary (CC-BY-SA-4.0) and PHOIBLE (CC-BY-SA-3.0) — keep their share-alike terms if you redistribute those fields.

Download the open data

A deterministic, reproducible snapshot of release 1.1.0. These are open data under the upstream licences — attribute accordingly. The atlas's own derived layer (the joins, tiers, provenance flags, coverage tables) is released CC-BY-4.0.

Prefer a single documented bundle? The README / data dictionary describes every file, column and licence. A CLDF package is not yet emitted; the CSV + GeoJSON + BibTeX + CITATION.cff set above is the interchange format for this release.

How to cite this atlas

Cite the atlas and the upstream datasets you rely on. The machine-readable CITATION.cff and the full sources.bib (upstream datasets + the scholarly references behind the movement theses) are in the downloads.

Suggested citation

Mob Translate project & Claude (2026). The Atlas of Australian Languages, data release 1.1.0. https://mobtranslate.com/atlas (accessed <date>).

BibTeX

@misc{atlas_australian_languages,
  author       = {{Mob Translate project and Claude (Anthropic)}},
  title        = {The Atlas of Australian Languages},
  year         = {2026},
  version      = {1.1.0},
  howpublished = {\url{https://mobtranslate.com/atlas}},
  note         = {Data release 1.1.0, build 2026-07-12. Derived layer CC-BY-4.0. Cite the upstream datasets too (see sources.bib).}
}

Method & reproducibility

A single committed build script, pnpm atlas:build-data, reads the canonical AIATSIS registry (1,029 languoids), a read-only snapshot of the live dictionary database, the Bouckaert 2018 phylogeography, and the Grambank/WALS/AUS typology plane, and joins them deterministically. A sibling step, pnpm atlas:build-exports, emits the download files above. The build is deterministic (the version and build stamp come from the manifest, never the wall clock), so a release is a git tag and re-running reproduces byte-identical artifacts.

Join keys. Everything is keyed on glottocode first, falling back to AUSTLANG code, then normalised name / DB slug, with a legacy-alias map so old links resolve.

Coordinate fallback chain. (1) Glottolog point; (2) AUSTLANG approximate location; (3) subgroup-centroid of sibling leaves, flagged derived; (4) if none, the language is not given a coordinate — it stays in the directory and CSV with coord_provenance=none. A coordinate is never fabricated to look complete.

What “genuine languoid” excludes. 49 Sign-language, Pidgin, Bookkeeping and mixed/artificial nodes are moved to a clearly-labelled appendix rather than counted as languages; Mi'gmaq is excluded as non-Australian; the Curr 1886-87 OCR wordlists are kept as a historical appendix, not as modern language identities.

Coverage, shown as fractions

Out of 980 genuine languoids (1,029 registry nodes minus 49 appendixed non-language nodes). Fractions are never rounded up to imply completeness.

located (with a real or approximate coordinate)
809 / 980
with a full classification chain
647 / 980
grammatically profiled (Grambank/WALS/AUS)
203 / 980
with a dated deep-time position
224 / 980
with a coded endangerment level
442 / 980
with a Glottolog code
656 / 980

Uncertainty & honesty policy

  • Coordinates carry provenance.

    Every point is flagged glottolog / austlang / derived_centroid / none. Approximate and derived points are labelled as such; unlocated languages are listed, never plotted at a fake point.

  • Autonyms are unverified candidates.

    Names drawn from alt-name lists (which conflate endonyms with spelling variants) are shown as candidates, never asserted as a community-confirmed autonym.

  • Grammar coverage is partial and honest.

    203 of 980 languoids are grammatically profiled; the rest are shown as not profiled, never as absence of a feature. The grammar lens reports grammar_recorded_agreement — agreement over jointly-recorded features only, not overall grammatical similarity and not genetic relatedness (always read n_joint).

  • Deep-time dates a language lineage, not a population.

    The Bouckaert 2018 ages (224 of 980 leaves dated) reconstruct the movement of language lineages across a continent already populated for ~65,000 years — not a peopling event. The deepest Pama-Nyungan nodes have weak posterior support and are shown with their 95% HPD, not a bare number. See the deep-time spread.

  • The movement theses preserve contradiction.

    The eight “why did the languages move?” theses are held side by side — each with mandatory evidence against and a contestation level. No single winner is declared. Read the thesis matrix.

Indigenous Data Sovereignty · CARE · ICIP

These languages belong to the First Peoples of this continent, who have spoken them on Country for tens of thousands of years and speak many of them today. This atlas aggregates open scholarly and catalogue records to help people find, cite and care for that knowledge. It is a scholarly aggregation — it is not a substitute for community authority, and it claims no community endorsement.

Names, autonyms and locations drawn from catalogues are shown with their uncertainty, not as settled fact. Rights-managed materials — AIATSIS and language-centre dictionaries — appear only as catalogue pointers; they are never scraped or reproduced. Nineteenth-century colonial wordlists are labelled as historical sources, not as community-approved lexicons. Communities are the final word on their own languages.

We follow the spirit of the CARE Principles for Indigenous Data Governance — Collective benefit, Authority to control, Responsibility, Ethics — alongside FAIR, and the Local Contexts / Traditional Knowledge (TK) & ICIP framework for Indigenous Cultural and Intellectual Property.

If any record here is wrong, sensitive, or should be withheld, please tell us — we will correct or remove it. Contact ajax@mobtranslate.com.

We pay our respects to Elders past and present, and to the language custodians and speakers whose knowledge this records.

Attribution & credits

With thanks to Glottolog, Grambank, WALS, PHOIBLE, D-PLACE, Phlorest, AIATSIS AUSTLANG, and to Bouckaert, Bowern & Atkinson, whose open work makes this possible; to E. M. Curr (1886–87) for the historical wordlists (colonial OCR, not community-approved lexicons); and, above all, to the language communities and custodians whose knowledge this records. Full references are in sources.bib.