Skip to main content

Toolkit for deriving morphosyntactic constituency spans from annotated planar structures

Project description

planars

A Python toolkit for deriving morphosyntactic constituency spans from annotated planar structures. Designed for cross-linguistic typological research.

Overview

Planar structures are ordered sequences of positions representing the morphosyntactic template of a language's verbal domain. Each position is filled by one or more elements (forms or form-types). Researchers annotate elements with diagnostic parameters, and this toolkit derives spans — ranges of positions identified as domains by various constituency tests.

Four span types are computed for each analysis:

Complete positions Partial positions
Strict (no gaps) strict complete strict partial
Loose (gaps allowed) loose complete loose partial

See codebook.yaml for definitions of all parameters, values, and terms.

This toolkit builds on the theoretical framework developed in:

Tallman, Adam J. R., Sandra Auderset, and Hiroto Uchihara (eds.). 2024. Constituency and convergence in the Americas. Topics in Phonological Diversity 1. Berlin: Language Science Press. doi:10.5281/zenodo.10559861

Requirements

  • Python 3.9+
  • pandas
  • gspread + google-auth + google-api-python-client (Google Sheets workflow only)
python -m venv .venv
.venv/bin/python -m pip install -r requirements.txt
.venv/bin/python -m ipykernel install --user --name planars --display-name planars

Workflow

1. Generate annotation forms

python generate_sheets.py   # creates sheets for new classes; skips existing ones
python generate_sheets.py --force  # regenerate all from scratch

Creates one Google Sheets file per analysis class with one tab per construction. Each tab has per-parameter dropdown validation and a free-text Comments column. Google Sheets is the definitive copy of annotation forms. On re-runs, only classes not yet in the Drive manifest are created.

Authentication uses OAuth2. On first run a browser window opens for authorization; the token is cached at ~/.config/gspread/authorized_user.json. OAuth credentials must be at ~/.config/planars/oauth_credentials.json (override with PLANARS_OAUTH_CREDENTIALS).

The manifest is stored on Google Drive as manifest_{lang_id}.json in the language folder. A local drive_config.json (gitignored) bootstraps the Drive lookup.

2. Annotate

Specialists fill in values in the shared Google Sheets. Keystone rows (v:verbroot) are pre-filled with NA and should not be changed.

3. Import

python import_sheets.py          # downloads filled sheets → TSVs in coded_data/
python import_sheets.py --force  # overwrite existing files

Skips existing files by default. If any validation warnings are found (blank cells, unexpected values), they are written to import_errors/{lang}_{timestamp}.txt as well as printed to the terminal.

4. Run analyses

From the repo root:

python -m planars ciscategorial     <path/to/filled.tsv>
python -m planars subspanrepetition <path/to/filled.tsv>
python -m planars noninterruption   <path/to/filled.tsv>
python -m planars stress            <path/to/filled.tsv>
python -m planars aspiration        <path/to/filled.tsv>

Maintaining sheets

python update_sheets.py           # dry run — show what would change
python update_sheets.py --apply   # add missing rows/trailing columns to existing sheets

python sync_params.py             # dry run — show param column changes needed
python sync_params.py --apply     # insert new param columns, update dropdown validation

Use update_sheets.py when new elements are added to the planar structure or a new trailing column (e.g. Comments) needs propagating. Use sync_params.py when diagnostics.tsv param columns change — it preserves existing annotations while inserting new columns before Comments.

5. Explore results interactively

source .venv/bin/activate
jupyter lab

Open notebooks/span_results.ipynb. Make sure the kernel in the top-right says planars (if not, go to Kernel → Change Kernel and select it). Run all cells with Run → Run All Cells. The notebook reads the filled TSVs directly and reports spans for all analyses, noting any positions with missing annotations.

Analyses

Analysis Parameters Spans derived
ciscategorial V-combines, N-combines, A-combines 4 (strict/loose × complete/partial)
subspanrepetition widescope_left, widescope_right, fillable_botheither_conjunct 20 (5 categories × 4)
noninterruption free, multiple 4 strict spans (2 domain types × complete/partial)
stress stressable, obligatory, independence, left-interaction, right-interaction 4 (provisional — qualification rule under review)
aspiration stressable, obligatory, independence, left-interaction, right-interaction 4 (provisional — qualification rule under review)

See codebook.yaml for qualification rules. Stress and aspiration entries are marked [NEEDS REVIEW].

Charting

planars.charts provides two functions for visualizing span results:

from planars.charts import collect_all_spans, domain_chart

df, keystone_pos, pos_to_name = collect_all_spans(repo_root)
fig = domain_chart(df, keystone_pos, pos_to_name)
fig.show()   # interactive Plotly figure
fig.write_image("domains.pdf")  # or save to file

collect_all_spans runs all analyses over all filled TSVs in coded_data/ and returns a DataFrame with columns Test_Labels, Analysis, Left_Edge, Right_Edge, Size. collect_all_spans_from_sheets(gc, manifest) does the same but reads directly from Google Sheets (used by Colab notebooks). domain_chart renders this as a horizontal segment chart with one row per span, colored by analysis type, with the keystone marked by a dotted line.

Colab notebooks

Two notebooks in notebooks/ support browser-only use without a local install:

  • sync_colab.ipynb — for non-technical collaborators: set DRIVE_FOLDER_PATH and run all cells to view the domain chart
  • span_results_colab.ipynb — for contributors: step-by-step cells for setup, manifest loading, and charting

diagnostics.tsv

Controls which analyses and constructions are generated for each language. Parameters default to y/n dropdowns; custom values use brace syntax:

stressable{y/n/both}, independence, left-interaction, right-interaction

Repository structure

planars/                        Core library
  io.py                         Shared TSV/Sheets loader
  spans.py                      Span computation functions
  ciscategorial.py              }
  subspanrepetition.py          }
  noninterruption.py            } Analysis modules
  stress.py                     }
  aspiration.py                 }
  charts.py                     Span collection and domain chart
  cli.py                        Command-line entry point
coded_data/{lang_id}/           Annotation data per language
  planar_input/                 Planar structure TSV + diagnostics.tsv
  {class_name}/                 Filled TSVs per analysis class
make_forms.py                   Planar structure and diagnostics utilities
generate_sheets.py              Create annotation forms in Google Drive
update_sheets.py                Add missing rows/trailing columns to existing sheets
sync_params.py                  Sync param columns when diagnostics.tsv changes
import_sheets.py                Download filled sheets to TSVs
restructure_sheets.py           Archive and regenerate sheets after structural changes
notebooks/                      Jupyter notebooks (local + Colab)
tests/snapshots/                Regression test baselines
codebook.yaml                   Parameter and term definitions

Regression testing

python generate_snapshots.py   # regenerate baselines
python check_snapshots.py      # verify output matches baselines

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

planars-0.1.0a2.tar.gz (17.8 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

planars-0.1.0a2-py3-none-any.whl (19.0 kB view details)

Uploaded Python 3

File details

Details for the file planars-0.1.0a2.tar.gz.

File metadata

  • Download URL: planars-0.1.0a2.tar.gz
  • Upload date:
  • Size: 17.8 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.13.1

File hashes

Hashes for planars-0.1.0a2.tar.gz
Algorithm Hash digest
SHA256 cedb93c4478325c214dadef1790d692df15f54144821151aa201b282f5f81eda
MD5 9682b7f02b6bdcfa7a195558b2ef9fd3
BLAKE2b-256 58cc7a3ae10b3d5f2864b95cea0ec87dfb51c6759ef8172da30db1955765240e

See more details on using hashes here.

File details

Details for the file planars-0.1.0a2-py3-none-any.whl.

File metadata

  • Download URL: planars-0.1.0a2-py3-none-any.whl
  • Upload date:
  • Size: 19.0 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.13.1

File hashes

Hashes for planars-0.1.0a2-py3-none-any.whl
Algorithm Hash digest
SHA256 e87eb39345ad9ecd8cb9903f90b19d08779830e1bdaec1adec559aded709cd7d
MD5 04624a6921a74ef2837ab5732d3740f6
BLAKE2b-256 c618ae0dacec9935c054f0793446711ef1e77c04fae0314831b5e9a3688cc6b1

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page