Body-Part Reflexives and Body-Zone Classification

From The Observatory
Jump to navigation Jump to search
The printable version is no longer supported and may have rendering errors. Please update your browser bookmarks and please use the default browser print function instead.

A Typological Survey with Substrate Implications

Preprint — Version 2 — July 2026

Jan Ritch-Frel and Ed Phillips

PREPRINT — Not peer reviewed

Abstract

This paper documents two typological phenomena and explores their possible connections. Track A: the grammaticalization of "head" nouns into reflexive pronouns (HEAD→SELF) clusters in the Caucasus at rates above the cross-linguistic background. A Bayesian generalized linear mixed model (brms; Bürkner 2017; formula: bpdr ~ caucasus + lat_c + lon_c + (1|Family); weakly informative priors; 4 chains × 2,000 draws) yields a Caucasus effect of OR = 10.30 (95% credible interval [1.82, 55.86]; 99.4% of posterior mass positive), based on 2,463 languages across 216 families from Grambank. The wide credible interval reflects genuine uncertainty in the effect size, though the model consistently favors a positive Caucasus effect. The same grammaticalization pathway is widespread across Afro-Asiatic, attested with independent lexical roots in five of six branches. Track B: Great Andamanese and Kusunda — both isolates in geographic refugia — share a structurally parallel body-zone noun classification architecture not attested in any other documented language family. A peer-reviewed reconstruction of Proto-Kusunda (Spendley 2024) independently identifies the parallel, describing the Kusunda system as "strongly reminiscent" of the GA morphology. The paper proposes an elaboration model and explores a substrate/dispersal interpretation as a hypothesis for future testing. The R analysis pipeline (brms/rstan) is provided for independent replication, cross-validated against an independent Python/bambi implementation.

Downloads

Contact the authors for the full preprint PDF and supplementary data (12 CSV tables + R analysis pipeline with testthat test suite).

Data Sources

  • Grambank v1.0.3 — Skirgård et al. (2023). "Grambank reveals the importance of genealogical constraints on linguistic diversity." Science Advances 9(16): eadg6175.
  • CLICS4 — Tjuka et al. (2025). Database of Cross-Linguistic Colexifications. Zenodo.
  • Lexibank v2.1 — Zenodo.
  • Evseeva & Salaberri (2018/2019). "Grammaticalization of nouns meaning 'head' into reflexive markers." Linguistic Typology 22(3): 385–435; Corrigendum 23(1): 255–262.
  • Spendley (2024). "Possessive prefixes in Proto-Kusunda." Himalayan Linguistics 23(1): 57–71.
  • Bürkner (2017). "brms: An R Package for Bayesian Multilevel Models Using Stan." Journal of Statistical Software 80(1): 1–28.

Version History

  • Version 2 (July 2026): Reframed as hypothesis-generating typological study. R/brms model (OR = 10.30). Co-authored with Ed Phillips. Removed deep-time inheritance headline claim. Added alternative explanations, falsifiability conditions.
  • Version 1 (June 2026): Initial preprint.

Template:Cc-by-nc-sa-4.0