Methods, sources & open data
How this atlas is built, and how to cite it
The Atlas of Australian Languages is assembled from open scholarly datasets and served as versioned static artifacts — never a live database query at request time. Every field on every profile carries its source, coverage is reported as honest fractions rather than rounded to imply completeness, and the whole joined dataset is downloadable below under the upstream licences. Data release 1.1.0 · build 2026-07-12.
Upstream datasets
Each source is used for a specific plane of the atlas. Versions and licences are exactly as released; where a versioned Zenodo DOI is best obtained from the dataset's own landing page we link the canonical page rather than assert an identifier we have not verified.
| Dataset | Version | Licence | Link / DOI | Used for |
|---|---|---|---|---|
| Glottolog | 5.3 CLDF | CC-BY-4.0 | Classification chains, coordinates, ISO/glottocodes, AES endangerment. | |
| Grambank | v1.0.3 CLDF | CC-BY-4.0 | Grammatical feature profiles (the /atlas/grammar lens + grammar-values.csv). | |
| WALS | 2020.4 CLDF | CC-BY-4.0 | Supplementary grammatical features baked into the feature tables. | |
| Australianist typology extension (AUS extension) | mobtranslate-pg typology plane | CC-BY-4.0 (derived) | Australia-specific features + recorded-agreement similarity neighbours. | |
| AIATSIS AUSTLANG | data.gov.au export | CC-BY-4.0 | Canonical AU language codes, approximate coordinates, alt-name lists. | |
| Bouckaert, Bowern & Atkinson 2018 (Phlorest) | phlorest CLDF | CC-BY-4.0 | The dated Pama-Nyungan phylogeographic tree (deep-time positions on /atlas/spread). | |
| PHOIBLE | referenced | CC-BY-SA-3.0 | Phonological inventories — referenced pointer only, not yet ingested. | |
| E. M. Curr, The Australian Race (1886-87) | archive.org OCR | Public Domain | 19th-century locality wordlists (historical appendix — colonial OCR, not community lexicons). | |
| English Wiktionary | kaikki.org export | CC-BY-SA-4.0 | Open lexical resources for some languoids. | |
| mobtranslate-pg (live DB) | read-only snapshot @ 2026-07-12T00:00:00Z | mixed (see per-source) | Live, community-curated dictionary word-counts (read-only build-time snapshot). |
CC-BY vs CC-BY-NC.The Bouckaert et al. 2018 phylogeny is used here via the Phlorest CLDF release (CC-BY-4.0). The same tree is also distributed through D-PLACE, parts of which carry CC-BY-NC. The atlas deliberately sources the CC-BY-4.0 Phlorest release so the deep-time layer is free of a non-commercial clause. Share-alike sources — Wiktionary (CC-BY-SA-4.0) and PHOIBLE (CC-BY-SA-3.0) — keep their share-alike terms if you redistribute those fields.
Download the open data
A deterministic, reproducible snapshot of release 1.1.0. These are open data under the upstream licences — attribute accordingly. The atlas's own derived layer (the joins, tiers, provenance flags, coverage tables) is released CC-BY-4.0.
- languages.csv
One row per genuine languoid: identity, family, coordinates + provenance, tier, endangerment, grammar/lexicon coverage, deep-time divergence and full classification chain.
CSV · 166 KB · 980 rows
- languages.geojson
GeoJSON FeatureCollection of the 809 located languages (Point geometry); the 171 unlocated languoids are honestly omitted, never given a fake point.
GEOJSON · 467 KB · 809 rows
- grammar-values.csv
Long-format grammatical feature codings (one row per language × feature). value/present/absent/unknown(?) states kept distinct — read the state column; unknown = not recorded.
CSV · 4.3 MB · 30,351 rows
- theses.json
The 8 "why did the languages move?" scholarly theses, contradiction-preserving — each with mechanism, proponents, evidence for AND against, hard facts, contestation level and full citations (with DOIs).
JSON · 126 KB · 8 rows
- sources.bib
BibTeX for the upstream datasets and the 48 scholarly references behind the movement theses (DOIs as published). Includes an entry for citing the atlas itself.
BIB · 27 KB · 56 rows
- CITATION.cff
Citation File Format metadata for the atlas itself (title, authors, version, licence, release date, upstream references) — machine-readable "cite this".
CFF · 3 KB
- README.md
Data dictionary: every file + column, the per-source licences, how to cite, and the honesty/uncertainty notes.
MD · 8 KB
Prefer a single documented bundle? The README / data dictionary describes every file, column and licence. A CLDF package is not yet emitted; the CSV + GeoJSON + BibTeX + CITATION.cff set above is the interchange format for this release.
How to cite this atlas
Cite the atlas and the upstream datasets you rely on. The machine-readable CITATION.cff and the full sources.bib (upstream datasets + the scholarly references behind the movement theses) are in the downloads.
Suggested citation
Mob Translate project & Claude (2026). The Atlas of Australian Languages, data release 1.1.0. https://mobtranslate.com/atlas (accessed <date>).BibTeX
@misc{atlas_australian_languages,
author = {{Mob Translate project and Claude (Anthropic)}},
title = {The Atlas of Australian Languages},
year = {2026},
version = {1.1.0},
howpublished = {\url{https://mobtranslate.com/atlas}},
note = {Data release 1.1.0, build 2026-07-12. Derived layer CC-BY-4.0. Cite the upstream datasets too (see sources.bib).}
}Method & reproducibility
A single committed build script, pnpm atlas:build-data, reads the canonical AIATSIS registry (1,029 languoids), a read-only snapshot of the live dictionary database, the Bouckaert 2018 phylogeography, and the Grambank/WALS/AUS typology plane, and joins them deterministically. A sibling step, pnpm atlas:build-exports, emits the download files above. The build is deterministic (the version and build stamp come from the manifest, never the wall clock), so a release is a git tag and re-running reproduces byte-identical artifacts.
Join keys. Everything is keyed on glottocode first, falling back to AUSTLANG code, then normalised name / DB slug, with a legacy-alias map so old links resolve.
Coordinate fallback chain. (1) Glottolog point; (2) AUSTLANG approximate location; (3) subgroup-centroid of sibling leaves, flagged derived; (4) if none, the language is not given a coordinate — it stays in the directory and CSV with coord_provenance=none. A coordinate is never fabricated to look complete.
What “genuine languoid” excludes. 49 Sign-language, Pidgin, Bookkeeping and mixed/artificial nodes are moved to a clearly-labelled appendix rather than counted as languages; Mi'gmaq is excluded as non-Australian; the Curr 1886-87 OCR wordlists are kept as a historical appendix, not as modern language identities.
Coverage, shown as fractions
Out of 980 genuine languoids (1,029 registry nodes minus 49 appendixed non-language nodes). Fractions are never rounded up to imply completeness.
- located (with a real or approximate coordinate)
- 809 / 980
- with a full classification chain
- 647 / 980
- grammatically profiled (Grambank/WALS/AUS)
- 203 / 980
- with a dated deep-time position
- 224 / 980
- with a coded endangerment level
- 442 / 980
- with a Glottolog code
- 656 / 980
Uncertainty & honesty policy
Coordinates carry provenance.
Every point is flagged
glottolog/austlang/derived_centroid/none. Approximate and derived points are labelled as such; unlocated languages are listed, never plotted at a fake point.Autonyms are unverified candidates.
Names drawn from alt-name lists (which conflate endonyms with spelling variants) are shown as candidates, never asserted as a community-confirmed autonym.
Grammar coverage is partial and honest.
203 of 980 languoids are grammatically profiled; the rest are shown as not profiled, never as absence of a feature. The grammar lens reports
grammar_recorded_agreement— agreement over jointly-recorded features only, not overall grammatical similarity and not genetic relatedness (always readn_joint).Deep-time dates a language lineage, not a population.
The Bouckaert 2018 ages (224 of 980 leaves dated) reconstruct the movement of language lineages across a continent already populated for ~65,000 years — not a peopling event. The deepest Pama-Nyungan nodes have weak posterior support and are shown with their 95% HPD, not a bare number. See the deep-time spread.
The movement theses preserve contradiction.
The eight “why did the languages move?” theses are held side by side — each with mandatory evidence against and a contestation level. No single winner is declared. Read the thesis matrix.
Indigenous Data Sovereignty · CARE · ICIP
These languages belong to the First Peoples of this continent, who have spoken them on Country for tens of thousands of years and speak many of them today. This atlas aggregates open scholarly and catalogue records to help people find, cite and care for that knowledge. It is a scholarly aggregation — it is not a substitute for community authority, and it claims no community endorsement.
Names, autonyms and locations drawn from catalogues are shown with their uncertainty, not as settled fact. Rights-managed materials — AIATSIS and language-centre dictionaries — appear only as catalogue pointers; they are never scraped or reproduced. Nineteenth-century colonial wordlists are labelled as historical sources, not as community-approved lexicons. Communities are the final word on their own languages.
We follow the spirit of the CARE Principles for Indigenous Data Governance — Collective benefit, Authority to control, Responsibility, Ethics — alongside FAIR, and the Local Contexts / Traditional Knowledge (TK) & ICIP framework for Indigenous Cultural and Intellectual Property.
If any record here is wrong, sensitive, or should be withheld, please tell us — we will correct or remove it. Contact ajax@mobtranslate.com.
We pay our respects to Elders past and present, and to the language custodians and speakers whose knowledge this records.
Attribution & credits
With thanks to Glottolog, Grambank, WALS, PHOIBLE, D-PLACE, Phlorest, AIATSIS AUSTLANG, and to Bouckaert, Bowern & Atkinson, whose open work makes this possible; to E. M. Curr (1886–87) for the historical wordlists (colonial OCR, not community-approved lexicons); and, above all, to the language communities and custodians whose knowledge this records. Full references are in sources.bib.